Voice filtering | Cariara Docs
Skip to content

DocsAttend

Voice filtering

Speaker verification — Cariara records your voiceprint once, then transcribes only the interviewer's voice during a live session.

Updated MarkdownOpen Live

Why this exists

When Cariara captures audio from your microphone (or a room mic that hears both of you), it picks up everything you say too. Without filtering, Sona would respond to your own answers — defeating the point. Voice filtering uses speaker verification to drop your voice from the transcription stream before it reaches Sona.

Enroll your voice

  1. Open the Live settings panel and click Enroll My Voice. (The button is also surfaced in the AI Companion panel and in older session UIs.)
  2. Talk in your normal voice for 15 seconds (about yourself, your work); that is enough — recording starts automatically and stops itself when finished.
  3. Cariara extracts a 256-dimension voiceprint embedding (a numeric fingerprint — no raw audio is stored long-term) and saves it to your account.
  4. Filtering is enabled automatically once enrollment finishes — you don't have to flip a separate switch. The button now reads Filter On; click it to toggle off.
Tip
Re-enroll if you have a cold or you've changed mics — embedding similarity drops when input characteristics change. Click Remove Enrollment, then enroll again with the new mic.
Note
On the “Room mic / any speaker” capture method, voice enrollment + filter are required — without them, Sona would treat your own voice as the interviewer.

How it works

For each audio chunk Cariara captures, the speaker service:

  1. Computes a speaker embedding for the chunk via resemblyzer.
  2. Compares cosine similarity to your enrolled voiceprint.
  3. For full-clip verification, similarity > 0.75 means the audio is yours — drop it. Otherwise it's the interviewer and gets transcribed.
  4. For longer clips with both speakers, Cariara runs sliding-window diarization at similarity threshold 0.55 to label which segments are yours vs the interviewer, then drops the user-labeled segments before transcription.

For panel interviews with multiple non-you speakers, all of them get bucketed as “interviewer” — Cariara doesn't separate panelists individually today.

When to disable it

  • Pure system-audio capture — if Cariara is only listening to system audio (Zoom loopback, tab share, virtual loopback), your mic isn't in the stream anyway, so filtering does nothing.
  • Solo practice — if you want Sona to respond to yourquestions for self-prep, turn the filter off.
  • Mismatch / cold — if Sona keeps missing the interviewer because your voiceprint is matching too aggressively, remove enrollment and re-enroll fresh.

Troubleshooting

  • “Sona is answering my voice” — your voiceprint matched too loosely. Re-enroll with the same mic you're using for the live call.
  • “Sona keeps missing the interviewer” — your voiceprint is dropping too much. Try removing enrollment and recording again (15 seconds, your normal voice).
  • Filter auto-disables mid-call — after a long run of consecutive drops, Cariara pauses the filter to keep the call moving. Re-enable from Settings if you still want it on.

Need help?