For AI Agents: This README is designed for agent onboarding. Read this first before making any changes.
A voice AI application built with Next.js that enables:
- Browser-based voice conversations with AI via OpenAI Realtime API (WebRTC)
- Phone call initiation via ElevenLabs + Twilio integration
This is a minimal markup version - function over form:
- Black background, white text (inverted high-contrast theme)
- Simple rectangular shapes (no rounded corners)
- Minimal styling - just enough to be functional
- No branding, polish, or aesthetic decisions yet
Why: This bare-bones version will be converted to a branded, polished app later. The skeleton must work first.
src/
├── app/ # Next.js App Router
│ ├── layout.tsx # Root layout
│ ├── page.tsx # Landing page
│ ├── globals.css # Minimal B&W styles
│ │
│ ├── voice/
│ │ └── page.tsx # Voice agent screen
│ │
│ └── api/ # Server-side routes (API keys stay here)
│ ├── token/route.ts # Returns ephemeral token + session config
│ └── call/route.ts # Initiates ElevenLabs phone call (placeholder)
│
├── components/ # UI Components
│ ├── VoiceButton.tsx # Start/stop with mic animations
│ ├── Transcript.tsx # Live conversation display
│ ├── EventLog.tsx # System events display
│ ├── PhoneCall.tsx # Phone input + call button
│ └── AudioVisualizer.tsx # Real-time frequency bars animation
│
└── lib/ # Core business logic
├── realtime.ts # OpenAI Realtime API (WebRTC client)
├── prompts.ts # System prompts and instructions
└── calling.ts # ElevenLabs phone call logic
| Layer | Purpose | Key Constraint |
|---|---|---|
app/api/* |
Server-side only | API keys cannot be in browser code |
lib/ |
Core logic | Framework-agnostic, all WebRTC/calling logic lives here |
components/ |
UI pieces | Each handles its own complexity (e.g., mic animations) |
app/voice/page.tsx |
Page composition | Glues components + lib together |
- Connection method: WebRTC (not WebSocket - that's deprecated for browsers)
- Auth flow:
- Client calls
/api/tokento get ephemeral session token - Client uses token to establish WebRTC connection
- Audio streams bidirectionally, events via DataChannel
- Client calls
All session settings are configured in /api/token/route.ts when creating the ephemeral token. This is the single source of truth.
| Setting | Value | Purpose |
|---|---|---|
model |
gpt-realtime |
Latest stable Realtime model |
voice |
marin |
AI voice selection |
instructions |
SYSTEM_PROMPT |
System prompt from lib/prompts.ts |
modalities |
['text', 'audio'] |
Enable both text and audio output |
input_audio_transcription.model |
gpt-4o-transcribe |
Transcription model for user speech |
turn_detection.type |
semantic_vad |
AI-powered turn detection |
input_audio_noise_reduction |
near_field |
Server-side noise suppression |
| input_audio_noise_reduction | near_field | Server-side noise suppression |
The client (lib/realtime.ts) does not send session.update - it only triggers the initial response after connection.
To avoid "double processing" artifacts (robotic/underwater voice), we split responsibilities:
- Browser (
lib/realtime.ts): HandlesechoCancellation(mandatory) andautoGainControl.noiseSuppressionis DISABLED here. - Server (OpenAI): Handles
noiseReduction(viainput_audio_noise_reduction: 'near_field').
Flow: Microphone → Browser Echo Cancellation → OpenAI Server Noise Reduction → Model
- Status: Implemented (Basic "Managed Service" integration)
- Endpoint:
https://api.elevenlabs.io/v1/convai/twilio/outbound-call - Flow: Client →
/api/call→ ElevenLabs API → Twilio → User's phone - Note: This uses the "Easy" mode where ElevenLabs manages the call leg. Live transcripts are not available in this mode.
OPENAI_API_KEY= # Required for Realtime API
ELEVENLABS_API_KEY= # Required for outbound calls
ELEVENLABS_AGENT_ID= # Required for outbound calls
ELEVENLABS_PHONE_NUMBER_ID= # Required for outbound calls (ID from ElevenLabs dashboard, starts with phnum_)
- Project structure scaffolded
- Placeholder files created
- Landing page (Grid layout, Black & White)
- Voice page layout
-
/api/tokenendpoint (OpenAI ephemeral token) -
lib/realtime.ts(WebRTC connection) - VoiceButton component (start/stop)
- Transcript component (live messages)
- EventLog component (system events)
-
/api/callendpoint (ElevenLabs) - PhoneCall component (with validation & loading states)
If at any point you determine that:
- The current architecture doesn't fit the requirements
- A different approach would be significantly better
- What you're being asked to implement conflicts with the existing structure
- The implementation is heading in a wrong direction
- There's a better way to do something
YOU MUST TELL THE USER IMMEDIATELY.
Don't silently adapt or work around issues. Speak up. The architecture can and should evolve as we learn more during implementation.
- Build incrementally: One feature at a time, test as you go
- Keep it simple: Hard-code values initially if needed
- Ask questions: If integration details are unclear, ask
- Markup first: Don't spend time on aesthetics yet
- TypeScript with explicit types
- Comments for complex logic only (not obvious code)
- Each file has a header comment explaining its purpose
npm install
npm run dev
# Opens at http://localhost:3000| Decision | Rationale | Date |
|---|---|---|
| WebRTC over WebSocket | OpenAI recommends WebRTC for browsers (WebSocket deprecated) | 2026-01-30 |
| Ephemeral token pattern | Can't expose API key in browser; server generates short-lived token | 2026-01-30 |
Single lib/ folder |
Simpler than lib/realtime/ + lib/elevenlabs/; can split later if needed |
2026-01-30 |
No types/ folder |
Types defined in the files that use them; can extract later if shared | 2026-01-30 |
| Prompts as Code | System prompts stored in lib/prompts.ts instead of markdown files for reliability |
2026-01-31 |
| Centered "Cockpit" Layout | Transitioning elements (Square User + Wide Agent visualizers) focuses user on the session | 2026-01-31 |
| Server-Side Session Config | All session settings (model, voice, instructions) moved to /api/token for single source of truth |
2026-01-31 |
| Lucide Icons | Switched from inline SVGs to Lucide React for cleaner code and maintainability | 2026-01-31 |
| ElevenLabs Managed Mode | Selected "Managed Service" integration for phone calls to avoid complex WebSocket relay infrastructure for MVP | 2026-02-05 |
| Unified Idle Layout | Kept PhoneCall component mounted during state transitions to preserve focus and dropdown interactions |
2026-02-05 |
| Phone Number ID | Switched from raw TWILIO_PHONE_NUMBER to ELEVENLABS_PHONE_NUMBER_ID as Managed Service requires the ID, not the number string |
2026-02-05 |
| Explicit Audio Constraints | Explicitly enabled echoCancellation and autoGainControl in WebRTC client; disabled noiseSuppression to use server-side alternative |
2026-02-05 |
| Server-Side Noise Cancellation | Enabled input_audio_noise_reduction: 'near_field' in OpenAI session config to replace browser's inferior suppression |
2026-02-05 |
Last updated: 2026-02-05