Original Reddit post

Hey everyone, I am currently training to work as an AI Audio & Speech QA Evaluator . To build up my portfolio, I am testing modern text-to-speech (TTS) models, finding subtle acoustic/pronunciation bugs, and documenting them in structured evaluation tables. I’d love your eyes and ears on my latest evaluation to see if my technical notes make sense and if my findings are accurate. 🎧 The Audio Sample Voice Engine: Fish Audio — “Sarah” model (Tagged: Conversational / Soft / Breathy ) Audio Link: [https://soundcloud.com/garethslocombe20/sarah-2026-10-09-02-53-to] Test Script: “To evaluate transient response, the low-pass filter was set to 500 Hertz to prevent lower-end distortion on micro-transducers.” 📋 Example Evaluation Table (What I found in this clip) Below is an example of the completed QA table I filled out after listening to this specific audio clip: 💬 How you can help me: Give it a listen: Do your ears hear the same issues I listed in the table (like the transducers mispronunciation? Review the table: Is this table format easy for a speech engineer or audio lead to read quickly? Terminology check: Are there any audio or phonetic terms I could phrase better? Like any general problems you notice with the speech? What categories should i use? Thanks a lot for any feedback or critiques! submitted by /u/axi5music

Originally posted by u/axi5music on r/ArtificialInteligence