Skip to main content
Custom recordings allow you to build hybrid voice agents that use your own pre-recorded audio for key parts of the conversation, while falling back to LLM-generated speech in your agent’s voice for dynamic responses. This gives you the best of both worlds – the emotional depth of real human speech and the flexibility of AI-generated dialogue.

Why use custom recordings?

  • Emotional variance – Real recordings carry natural intonation and emotion that TTS cannot fully replicate.
  • Lower latency – Playing a pre-recorded clip is faster than synthesizing text at runtime.

Step 1: Upload recordings

Navigate to the Recordings page in the Bananaflow dashboard. Recordings are shared across all agents in your organization. You can either upload pre-recorded audio files or record directly in the browser. For each recording:
  1. Click Upload Recording.
  2. Choose an audio file or click Record to record in the browser.
  3. Verify the transcription is correct – edit it if needed.
  4. Click Upload.
You can rename a recording’s ID at any time by clicking the edit icon next to it in the recordings list.

Step 2: Build the workflow

Open your agent’s workflow and write the conversation flow in natural language. To insert a recording, type @ in the prompt editor – this will show a list of all available recordings in your organization. For any user question that falls outside your recordings, the agent automatically generates a dynamic response using the LLM, which is then spoken in the agent’s voice.

Tips for best results

  • Record in a quiet environment to improve audio quality.
  • Keep recordings concise – short, focused clips work best for specific conversation moments.
  • Review call recordings after testing to identify where the transition between pre-recorded and dynamic audio can be improved.