Build a Voice Agent in Claude: Step-by-Step Guide | desplega.ai
Learn how to build a voice agent using Claude's API. Connect speech-to-text, Claude, and text-to-speech to create a fully functional voice assistant.
Prerequisites
- Node.js 18+ installed on your machine
- An Anthropic API key with access to Claude models
- An OpenAI API key for Whisper STT and TTS (or a compatible alternative)
- A working microphone and speakers accessible to your OS
- Basic familiarity with TypeScript / Node.js async patterns
Voice agents have moved from novelty demos to production tools — customer support bots, kiosk assistants, accessibility helpers, and hands-free dev tooling all rely on the same three-stage pipeline: capture audio, reason over it, speak back. Claude is an excellent reasoning core for this pipeline because it follows multi-turn instructions reliably and produces natural, conversational replies that sound good when spoken aloud. In this guide we'll wire Claude up to OpenAI Whisper for speech-to-text and OpenAI TTS for playback, but every component is swappable. If you're already using Claude inside your editor, you may also want to read our companion guide on setting up MCP with Claude Code to give your voice agent access to tools and project context.
Set Up Your Project and Install Dependencies
Create a new project directory and install the required libraries for audio capture, the Anthropic SDK, and a text-to-speech provider. We'll usenode-record-lpcm16for microphone capture and speakerfor raw PCM playback — both are battle-tested on macOS and Linux.
mkdir claude-voice-agent && cd claude-voice-agent
npm init -y
npm install @anthropic-ai/sdk openai dotenv node-record-lpcm16 speakerConfigure API Keys in Your Environment
Create a .env file at your project root and add your Anthropic API key plus your chosen speech provider key so credentials stay out of source code. Add .env to your .gitignore immediately — leaked keys are by far the most common rookie mistake when shipping voice agents.
# .env
ANTHROPIC_API_KEY=sk-ant-...
OPENAI_API_KEY=sk-... # used for Whisper STT + TTSCapture Audio and Transcribe with Speech-to-Text
Record microphone input as a WAV buffer and send it to OpenAI Whisper (or any STT provider) to get a plain-text transcript of what the user said. We use a fixed 5-second window for simplicity; in production you'll want voice activity detection (VAD) so the agent knows when the user has finished speaking instead of cutting them off.
import Anthropic from '@anthropic-ai/sdk';
import OpenAI from 'openai';
import recorder from 'node-record-lpcm16';
import { createWriteStream } from 'fs';
import fs from 'fs';
const openai = new OpenAI();
async function recordAndTranscribe(durationMs = 5000): Promise<string> {
const file = '/tmp/input.wav';
await new Promise<void>((resolve) => {
const rec = recorder.record({ sampleRate: 16000 });
rec.stream().pipe(createWriteStream(file));
setTimeout(() => { rec.stop(); resolve(); }, durationMs);
});
const transcription = await openai.audio.transcriptions.create({
file: fs.createReadStream(file),
model: 'whisper-1',
});
return transcription.text;
}Send the Transcript to Claude and Get a Response
Pass the transcribed text as a user message to the Claude API, maintaining a conversation history array so the agent preserves context across turns. The system prompt is critical here — instruct Claude to keep replies short and conversational, because long paragraphs sound robotic when spoken and add latency to every turn.
const anthropic = new Anthropic();
const history: { role: 'user' | 'assistant'; content: string }[] = [];
async function askClaude(userText: string): Promise<string> {
history.push({ role: 'user', content: userText });
const response = await anthropic.messages.create({
model: 'claude-opus-4-5',
max_tokens: 1024,
system: 'You are a helpful voice assistant. Keep responses concise and conversational.',
messages: history,
});
const assistantText = response.content[0].type === 'text' ? response.content[0].text : '';
history.push({ role: 'assistant', content: assistantText });
return assistantText;
}Convert Claude's Response to Speech and Play It
Feed Claude's text reply into a TTS API to get an audio buffer, then stream it to your system speaker so the user hears the response. For lower latency you can switch to a streaming TTS endpoint and start playback before the full audio arrives — the perceived responsiveness improvement is huge.
import Speaker from 'speaker';
import { Readable } from 'stream';
async function speak(text: string): Promise<void> {
const mp3 = await openai.audio.speech.create({
model: 'tts-1',
voice: 'alloy',
input: text,
});
const buffer = Buffer.from(await mp3.arrayBuffer());
const readable = Readable.from(buffer);
const speaker = new Speaker({ channels: 1, bitDepth: 16, sampleRate: 24000 });
readable.pipe(speaker);
await new Promise((resolve) => speaker.on('close', resolve));
}Wire the Pipeline into a Continuous Conversation Loop
Combine recording, transcription, Claude inference, and speech playback into a loop that keeps the voice agent listening until the user says a wake word to exit. The skip on empty transcripts prevents Claude from hallucinating responses to silence — a small detail that makes the agent feel much more polished in real-world use.
async function runVoiceAgent() {
console.log('Voice agent started. Say "goodbye" to exit.');
while (true) {
console.log('Listening...');
const userText = await recordAndTranscribe(5000);
if (!userText.trim()) continue;
console.log('You:', userText);
if (userText.toLowerCase().includes('goodbye')) {
await speak('Goodbye! Have a great day.');
break;
}
const reply = await askClaude(userText);
console.log('Claude:', reply);
await speak(reply);
}
}
runVoiceAgent().catch(console.error);Test and Validate Your Voice Agent End-to-End
Run the agent locally, speak a few prompts, and verify that audio is captured, Claude responds correctly, and the TTS output is intelligible before adding it to your CI pipeline. Once the happy path works, write automated tests that replay canned audio fixtures through the pipeline so regressions in prompt changes, model upgrades, or dependency bumps fail loudly before they reach users.
# Run the agent
npx ts-node index.ts
# Optional: run automated E2E tests with desplega.ai
# to validate the full voice pipeline on every deployWith these seven steps you have a working voice agent powered by Claude. From here, the natural extensions are tool use (let Claude call functions during a conversation), barge-in support (let users interrupt the assistant mid-sentence), and streaming both Claude tokens and TTS audio for sub-second latency. If you want a deeper look at how to ship and monitor AI features end to end, check the desplega.ai blog for production playbooks.
Related Guides
Issues
Track and manage test failures and issues effectively. Learn how to identify, categorize, and resolve testing issues in your workflow.
Vibe QA Extension
Power your vibe coding workflow with AI-powered QA testing. The Desplega.ai Vibe QA Extension integrates seamlessly with Lovable, bringing enterprise-grade testing directly into your development workflow.
Personas
Test your application with different user personas to ensure comprehensive coverage. Learn how to create and use personas in your testing strategy.