I always associated the uncanny valley with faces and robots but I’m wondering if there’s a version of it for speech too. Some AI voices are obviously synthetic and you kind of accept them for what they are. But once a voice gets extremely close to human, the little things that are still off start standing out more. The timing is too clean, every sentence lands perfectly, nobody hesitates or corrects themselves. I watched this roundtable about speech models where they argued that perfect speech might actually be the wrong goal and it got me thinking about this. Can AI speech eventually get past that uncanny valley? submitted by /u/FieldMedical7537
Originally posted by u/FieldMedical7537 on r/ArtificialInteligence
You must log in or # to comment.
