After exploring the Sesame (.com) demo, I was both impressed and intrigued by its approach to voice synthesis. Sesame’s design strives for genuine “voice presence” through emotional intelligence, natural timing, and contextual awareness. Yet there’s a catch: while the output is remarkably lifelike, the system only processes a text transcript of what I say, missing the subtle inflections, laughter, and pauses that truly convey emotion. Imagine the next step—a system that fully listens to your voice and captures every nuance!
Watch the attached video extract for a glimpse into my conversation with Sesame. And if you’re curious, let me know in the comments if you’d like a full podcast episode where I interview AIs about this and other groundbreaking ideas.














