The transcript looked fine. The conversation was terrible.
The only way to test voice AI is to go through it. Out loud. Every time. It is painstaking in a way most software development is not, and completely worth it.
Building voice AI is nothing like building software. You can read a prompt response and think it sounds fine. You can even think it sounds good. Then you hear it spoken and realise immediately that it doesn't work: the rhythm is wrong, the pause is too short, the question lands like an interrogation instead of a conversation. The only way to test voice is to go through it. Out loud. From the beginning. Every time. It is painstaking in a way that most software development is not. But here is the thing that surprised me. The more we iterated on Blu (our AI discovery specialist) the more we realised it was doing something we hadn't fully anticipated. It wasn't just saving us time on intake calls. It was actually helping develop the ideas coming through the door. Because Blu is not afraid to ask the questions that need to be asked. Has anyone actually spoken to the people who would use this day to day? What needs to be true for this to work that nobody has confirmed yet? What does this break if it actually succeeds? In a first meeting with a real person, those questions get softened or avoided because there is a dynamic to manage. Blu has no such hesitation. It just asks. And people answer honestly because the conversation feels safe rather than pressured. Getting it to that point, though, is where the real work is. Writing the question is the easy bit. Making it land the right way, at the right moment, without feeling like someone reading from a script, takes more iteration than you'd think. And you cannot do it by reading. You have to hear it. Dozens of times. Until it stops feeling like a system asking and starts feeling like someone who actually wants to know. We have had briefs come through that were sharper and more considered than anything produced from a traditional intake call. Not because the person arrived better prepared. Because the right questions were asked at the right moment without anyone worrying about how they landed. That was not what we built it for. It was a happy consequence of building it properly. The slog is worth it. Eventually. --- If you want to see what it feels like from the other side, talk to Blu and find out what questions come up when nobody is worried about how they land. We wrote about the full story of building Blu in our case study, eighteen months of voice AI work that finally pointed inward.