# AssemblyAI Documentation > AssemblyAI provides production-ready Voice AI models for speech-to-text, speaker detection, sentiment analysis, and more. Build with pre-recorded or streaming audio using REST APIs and WebSockets. Generated from the local Mintlify documentation source. Content inside `` is excluded, and content inside `` is included. ## Contents - [AssemblyAI Documentation](#assemblyai-documentation-docs-index-mdx) - [Build with AI coding agents](#build-with-ai-coding-agents-docs-coding-agent-prompts-mdx) - [Models](#models-docs-getting-started-models-mdx) - [Evaluations](#evaluations-docs-evaluations-mdx) - [Account Management](#account-management-docs-account-management-mdx) - [Single Sign-On (SSO)](#single-sign-on-sso-docs-sso-mdx) - [Billing and Pricing](#billing-and-pricing-docs-billing-and-pricing-mdx) - [Introducing Universal-3.5 Pro](#introducing-universal-3-5-pro-docs-getting-started-universal-3-5-pro-mdx) - [End-to-end examples](#end-to-end-examples-docs-getting-started-end-to-end-examples-mdx) - [Meeting notetaker](#meeting-notetaker-docs-getting-started-end-to-end-examples-meeting-notetaker-mdx) - [Sales call intelligence](#sales-call-intelligence-docs-getting-started-end-to-end-examples-sales-call-intelligence-mdx) - [Medical scribe](#medical-scribe-docs-getting-started-end-to-end-examples-medical-scribe-mdx) - [Content repurposing](#content-repurposing-docs-getting-started-end-to-end-examples-content-repurposing-mdx) - [Real-time meeting assistant](#real-time-meeting-assistant-docs-getting-started-end-to-end-examples-real-time-meeting-assistant-mdx) - [Real-time live captioner](#real-time-live-captioner-docs-getting-started-end-to-end-examples-real-time-live-captioner-mdx) - [Best Practices for building Meeting Notetakers](#best-practices-for-building-meeting-notetakers-docs-meeting-notetaker-best-practices-mdx) - [Build a medical scribe](#build-a-medical-scribe-docs-medical-scribe-best-practices-mdx) - [Build a Post-Visit Medical Scribe](#build-a-post-visit-medical-scribe-docs-medical-scribe-best-practices-medical-scribe-post-visit-mdx) - [Build a Real-Time Medical Scribe](#build-a-real-time-medical-scribe-docs-medical-scribe-best-practices-medical-scribe-real-time-mdx) - [Best Practices for Building Voice Agents](#best-practices-for-building-voice-agents-docs-voice-agents-best-practices-mdx) - [Best Practices for building Contact Center Applications](#best-practices-for-building-contact-center-applications-docs-contact-center-best-practices-mdx) - [Integrations](#integrations-docs-integrations-mdx) - [Universal 3.5 Pro Realtime on LiveKit](#universal-3-5-pro-realtime-on-livekit-docs-voice-agents-livekit-universal-3-5-pro-mdx) - [Universal 3.5 Pro Realtime on Pipecat](#universal-3-5-pro-realtime-on-pipecat-docs-voice-agents-pipecat-universal-3-5-pro-mdx) - [Zapier Integration with AssemblyAI](#zapier-integration-with-assemblyai-docs-integrations-zapier-mdx) - [Integrate Make with AssemblyAI](#integrate-make-with-assemblyai-docs-integrations-make-mdx) - [The AssemblyAI n8n Integration](#the-assemblyai-n8n-integration-docs-integrations-n-8-n-mdx) - [The Postman collection for the AssemblyAI API](#the-postman-collection-for-the-assemblyai-api-docs-integrations-postman-mdx) - [Build a Zoom Real-time transcription bot with Recall.ai](#build-a-zoom-real-time-transcription-bot-with-recall-ai-docs-integrations-recall-mdx) - [Transcribe Your Zoom Meetings](#transcribe-your-zoom-meetings-docs-integrations-zoom-rtms-mdx) - [Integrate Telnyx with AssemblyAI](#integrate-telnyx-with-assemblyai-docs-integrations-telnyx-mdx) - [Integrate Twilio with AssemblyAI](#integrate-twilio-with-assemblyai-docs-integrations-twilio-mdx) - [Transcribe Your Amazon Connect Recordings](#transcribe-your-amazon-connect-recordings-docs-integrations-amazon-connect-mdx) - [Transcribe Genesys Cloud Recordings with AssemblyAI](#transcribe-genesys-cloud-recordings-with-assemblyai-docs-integrations-genesys-cloud-mdx) - [\U0001F99C️\U0001F517 LangChain Integration with AssemblyAI](#u0001f99c-u0001f517-langchain-integration-with-assemblyai-docs-integrations-langchain-mdx) - [\U0001F99C️\U0001F517 LangChain Python Integration with AssemblyAI](#u0001f99c-u0001f517-langchain-python-integration-with-assemblyai-docs-integrations-langchain-python-mdx) - [\U0001F99C️\U0001F517 LangChain JavaScript Integration with AssemblyAI](#u0001f99c-u0001f517-langchain-javascript-integration-with-assemblyai-docs-integrations-langchain-js-mdx) - [▲ Vercel AI SDK Integration with AssemblyAI](#vercel-ai-sdk-integration-with-assemblyai-docs-integrations-vercel-ai-sdk-mdx) - [Integrate Power Automate with AssemblyAI](#integrate-power-automate-with-assemblyai-docs-integrations-power-automate-mdx) - [Semantic Kernel Integration for AssemblyAI](#semantic-kernel-integration-for-assemblyai-docs-integrations-semantic-kernel-mdx) - [Integrate Activepieces with AssemblyAI](#integrate-activepieces-with-assemblyai-docs-integrations-activepieces-mdx) - [Haystack Integration for AssemblyAI](#haystack-integration-for-assemblyai-docs-integrations-haystack-mdx) - [Cloudflare](#cloudflare-docs-nav-links-cloudflare-mdx) - [Relay.app](#relay-app-docs-nav-links-relay-app-mdx) - [Bubble by Knowcode](#bubble-by-knowcode-docs-nav-links-bubble-by-knowcode-mdx) - [Pipedream](#pipedream-docs-nav-links-pipedream-mdx) - [Drupal](#drupal-docs-nav-links-drupal-mdx) - [Trust center](#trust-center-docs-nav-links-trust-center-mdx) - [Security overview](#security-overview-docs-nav-links-security-overview-mdx) - [Data Controls](#data-controls-docs-data-controls-mdx) - [Data retention and model training](#data-retention-and-model-training-docs-data-retention-and-model-training-mdx) - [API Reference](#api-reference-docs-api-reference-overview-mdx) - [Quickstart](#quickstart-docs-pre-recorded-audio-getting-started-transcribe-an-audio-file-mdx) - [Model selection](#model-selection-docs-pre-recorded-audio-select-the-speech-model-mdx) - [Prompting and Keyterms](#prompting-and-keyterms-docs-pre-recorded-audio-universal-3-5-pro-prompting-mdx) - [Medical Mode](#medical-mode-docs-pre-recorded-audio-medical-mode-mdx) - [Speaker Diarization](#speaker-diarization-docs-pre-recorded-audio-label-speakers-mdx) - [Multichannel Transcription](#multichannel-transcription-docs-pre-recorded-audio-transcribe-multiple-audio-channels-mdx) - [Code Switching](#code-switching-docs-pre-recorded-audio-code-switching-mdx) - [Automatic Language Detection](#automatic-language-detection-docs-pre-recorded-audio-language-detection-mdx) - [Set Language Manually](#set-language-manually-docs-pre-recorded-audio-set-language-manually-mdx) - [Supported Languages](#supported-languages-docs-pre-recorded-audio-supported-languages-mdx) - [Custom Spelling](#custom-spelling-docs-pre-recorded-audio-correct-spelling-of-terms-mdx) - [Filler Words](#filler-words-docs-pre-recorded-audio-include-filler-words-mdx) - [Word Search](#word-search-docs-pre-recorded-audio-search-for-words-in-transcript-mdx) - [Set the Start and End of the Transcript](#set-the-start-and-end-of-the-transcript-docs-pre-recorded-audio-set-the-start-and-end-of-the-transcript-mdx) - [Transcript Status](#transcript-status-docs-pre-recorded-audio-check-transcript-status-mdx) - [Transcript export options](#transcript-export-options-docs-pre-recorded-audio-export-transcripts-as-srt-vtt-or-text-mdx) - [Delete Transcripts](#delete-transcripts-docs-pre-recorded-audio-delete-transcripts-mdx) - [Upload a media file](#upload-a-media-file-docs-pre-recorded-audio-api-reference-files-upload-mdx) - [Submit a transcript](#submit-a-transcript-docs-pre-recorded-audio-api-reference-transcripts-submit-mdx) - [Get a transcript](#get-a-transcript-docs-pre-recorded-audio-api-reference-transcripts-get-mdx) - [Get transcript sentences](#get-transcript-sentences-docs-pre-recorded-audio-api-reference-transcripts-get-sentences-mdx) - [Get transcript paragraphs](#get-transcript-paragraphs-docs-pre-recorded-audio-api-reference-transcripts-get-paragraphs-mdx) - [Get subtitles](#get-subtitles-docs-pre-recorded-audio-api-reference-transcripts-get-subtitles-mdx) - [Get redacted audio](#get-redacted-audio-docs-pre-recorded-audio-api-reference-transcripts-get-redacted-audio-mdx) - [Word search](#word-search-docs-pre-recorded-audio-api-reference-transcripts-word-search-mdx) - [List transcripts](#list-transcripts-docs-pre-recorded-audio-api-reference-transcripts-list-mdx) - [Delete a transcript](#delete-a-transcript-docs-pre-recorded-audio-api-reference-transcripts-delete-mdx) - [Webhooks for pre-recorded audio](#webhooks-for-pre-recorded-audio-docs-pre-recorded-audio-webhooks-mdx) - [Benchmarks](#benchmarks-docs-pre-recorded-audio-benchmarks-mdx) - [Evaluating Pre-recorded STT models](#evaluating-pre-recorded-stt-models-docs-pre-recorded-audio-evaluations-mdx) - [Cloud Endpoints and Data Residency](#cloud-endpoints-and-data-residency-docs-pre-recorded-audio-select-the-region-mdx) - [Rate limits](#rate-limits-docs-pre-recorded-audio-rate-limits-mdx) - [Migration guide: Deepgram to AssemblyAI](#migration-guide-deepgram-to-assemblyai-docs-pre-recorded-audio-migration-guides-dg-to-aai-mdx) - [Migration guide: OpenAI to AssemblyAI](#migration-guide-openai-to-assemblyai-docs-pre-recorded-audio-migration-guides-oai-to-aai-mdx) - [Migration guide: AWS Transcribe to AssemblyAI](#migration-guide-aws-transcribe-to-assemblyai-docs-pre-recorded-audio-migration-guides-aws-to-aai-mdx) - [Migration guide: Google Speech-to-Text to AssemblyAI](#migration-guide-google-speech-to-text-to-assemblyai-docs-pre-recorded-audio-migration-guides-google-to-aai-mdx) - [Migration guide: Gladia to AssemblyAI](#migration-guide-gladia-to-assemblyai-docs-pre-recorded-audio-migration-guides-gladia-to-aai-mdx) - [Best Practices for building Meeting Notetakers](#best-practices-for-building-meeting-notetakers-docs-meeting-notetaker-best-practices-mdx) - [Build a medical scribe](#build-a-medical-scribe-docs-medical-scribe-best-practices-mdx) - [Best Practices for building Contact Center Applications](#best-practices-for-building-contact-center-applications-docs-contact-center-best-practices-mdx) - [Get YouTube Video Transcripts with yt-dlp](#get-youtube-video-transcripts-with-yt-dlp-docs-pre-recorded-audio-guides-transcribe-youtube-videos-mdx) - [Build a UI for Transcription with Gradio and Python](#build-a-ui-for-transcription-with-gradio-and-python-docs-pre-recorded-audio-guides-gradio-frontend-mdx) - [Detect Low Confidence Words in a Transcript](#detect-low-confidence-words-in-a-transcript-docs-pre-recorded-audio-guides-detecting-low-confidence-words-mdx) - [Running Bulk Transcription and Load Tests at Scale](#running-bulk-transcription-and-load-tests-at-scale-docs-pre-recorded-audio-guides-bulk-transcription-and-load-tests-at-scale-mdx) - [Transcribe Multiple Files Simultaneously Using the JavaScript SDK](#transcribe-multiple-files-simultaneously-using-the-javascript-sdk-docs-pre-recorded-audio-guides-sdk-node-batch-mdx) - [Transcribe Multiple Files Simultaneously Using the Python SDK](#transcribe-multiple-files-simultaneously-using-the-python-sdk-docs-pre-recorded-audio-guides-batch-transcription-mdx) - [Transcribe from an S3 Bucket](#transcribe-from-an-s3-bucket-docs-pre-recorded-audio-guides-transcribe-from-s3-mdx) - [Transcribe Google Drive Files](#transcribe-google-drive-files-docs-pre-recorded-audio-guides-transcribing-google-drive-file-mdx) - [Transcribe GitHub Files](#transcribe-github-files-docs-pre-recorded-audio-guides-transcribing-github-files-mdx) - [Iterate over Speaker Labels with Make.com](#iterate-over-speaker-labels-with-make-com-docs-pre-recorded-audio-guides-make-speaker-labels-mdx) - [Calculate the Talk / Listen Ratio of Speakers](#calculate-the-talk-listen-ratio-of-speakers-docs-pre-recorded-audio-guides-talk-listen-ratio-mdx) - [Plot A Speaker Timeline with Matplotlib](#plot-a-speaker-timeline-with-matplotlib-docs-pre-recorded-audio-guides-speaker-timeline-mdx) - [Generate Custom Speaker Labels with Pyannote](#generate-custom-speaker-labels-with-pyannote-docs-pre-recorded-audio-guides-use-assemblyai-with-pyannote-to-generate-custom-speaker-labels-mdx) - [Use Speaker Diarization with Async Chunking](#use-speaker-diarization-with-async-chunking-docs-pre-recorded-audio-guides-speaker-diarization-with-async-chunking-mdx) - [Setup A Speaker Identification System using Pinecone & Nvidia TitaNet](#setup-a-speaker-identification-system-using-pinecone-nvidia-titanet-docs-pre-recorded-audio-guides-titanet-speaker-identification-mdx) - [Use Automatic Language Detection as a Separate Step From Transcription](#use-automatic-language-detection-as-a-separate-step-from-transcription-docs-pre-recorded-audio-guides-automatic-language-detection-separate-mdx) - [Route to Default Language if Language Confidence is Low](#route-to-default-language-if-language-confidence-is-low-docs-pre-recorded-audio-guides-automatic-language-detection-route-default-language-mdx) - [Create Custom Length Subtitles](#create-custom-length-subtitles-docs-pre-recorded-audio-guides-subtitle-creation-by-word-count-mdx) - [Create Subtitles with Speaker Labels](#create-subtitles-with-speaker-labels-docs-pre-recorded-audio-guides-speaker-labelled-subtitles-mdx) - [Generate Subtitles for Videos](#generate-subtitles-for-videos-docs-pre-recorded-audio-guides-subtitles-mdx) - [Translate an AssemblyAI Subtitle Transcript](#translate-an-assemblyai-subtitle-transcript-docs-pre-recorded-audio-guides-translate-subtitles-mdx) - [Schedule a DELETE request with AssemblyAI and EasyCron](#schedule-a-delete-request-with-assemblyai-and-easycron-docs-pre-recorded-audio-guides-schedule-delete-mdx) - [Troubleshoot Common Errors](#troubleshoot-common-errors-docs-pre-recorded-audio-guides-common-errors-and-solutions-mdx) - [Implement Retry Server Error Logic](#implement-retry-server-error-logic-docs-pre-recorded-audio-guides-retry-server-error-mdx) - [Implement Retry Upload Error Logic](#implement-retry-upload-error-logic-docs-pre-recorded-audio-guides-retry-upload-error-mdx) - [Identify Duplicate Dual Channel Files](#identify-duplicate-dual-channel-files-docs-pre-recorded-audio-guides-identify-duplicate-channels-mdx) - [Correct Audio Duration Discrepancies with Multi-Tool Validation and Transcoding](#correct-audio-duration-discrepancies-with-multi-tool-validation-and-transcoding-docs-pre-recorded-audio-guides-audio-duration-fix-mdx) - [Audio File Downsampling Recommendations and Best Practices](#audio-file-downsampling-recommendations-and-best-practices-docs-pre-recorded-audio-guides-downsampling-mdx) - [Translate an AssemblyAI Subtitle Transcript](#translate-an-assemblyai-subtitle-transcript-docs-pre-recorded-audio-guides-translate-subtitles-mdx) - [Translate AssemblyAI Transcripts Into Other Languages Using Commercial Models](#translate-assemblyai-transcripts-into-other-languages-using-commercial-models-docs-pre-recorded-audio-guides-translate-transcripts-mdx) - [Transform Chinese transcripts into Simplified or Traditional Text](#transform-chinese-transcripts-into-simplified-or-traditional-text-docs-pre-recorded-audio-guides-traditional-simplified-chinese-mdx) - [Do More With Our SDKs](#do-more-with-our-sdks-docs-pre-recorded-audio-guides-do-more-with-sdk-mdx) - [Quickstart](#quickstart-docs-streaming-getting-started-transcribe-streaming-audio-mdx) - [Model selection](#model-selection-docs-streaming-select-the-speech-model-mdx) - [Optimizing Accuracy and Latency](#optimizing-accuracy-and-latency-docs-streaming-getting-started-optimizing-accuracy-and-latency-mdx) - [Turn Detection](#turn-detection-docs-streaming-turn-detection-mdx) - [Prompting and Keyterms](#prompting-and-keyterms-docs-streaming-prompting-and-keyterms-mdx) - [Conversation Context](#conversation-context-docs-streaming-universal-3-5-pro-context-carryover-mdx) - [Multilingual Transcription](#multilingual-transcription-docs-streaming-multilingual-transcription-mdx) - [Medical Mode](#medical-mode-docs-streaming-medical-mode-mdx) - [Voice Focus](#voice-focus-docs-streaming-voice-focus-mdx) - [Streaming Diarization and Multichannel](#streaming-diarization-and-multichannel-docs-streaming-label-speakers-and-separate-channels-mdx) - [PII Redaction](#pii-redaction-docs-streaming-pii-redaction-mdx) - [Filter profanity](#filter-profanity-docs-streaming-filter-profanity-from-transcripts-mdx) - [Message Sequence](#message-sequence-docs-streaming-message-sequence-mdx) - [Streaming WebSocket API](#streaming-websocket-api-docs-streaming-api-spec-streaming-websocket-mdx) - [Generate streaming token](#generate-streaming-token-docs-streaming-api-spec-generate-streaming-token-mdx) - [Authenticate with a temporary token](#authenticate-with-a-temporary-token-docs-streaming-authenticate-with-a-temporary-token-mdx) - [Webhooks for streaming speech-to-text](#webhooks-for-streaming-speech-to-text-docs-streaming-webhooks-mdx) - [Updating Configuration Mid-Stream](#updating-configuration-mid-stream-docs-streaming-updating-configuration-mid-stream-mdx) - [Common session errors and closures](#common-session-errors-and-closures-docs-streaming-common-session-errors-and-closures-mdx) - [Benchmarks](#benchmarks-docs-streaming-benchmarks-mdx) - [Evaluating Real-time STT models for Voice Agents](#evaluating-real-time-stt-models-for-voice-agents-docs-streaming-evaluations-voice-agents-mdx) - [Streaming Endpoints and Data Zones](#streaming-endpoints-and-data-zones-docs-streaming-endpoints-and-data-zones-mdx) - [Rate limits](#rate-limits-docs-streaming-rate-limits-mdx) - [Self-Hosted Streaming](#self-hosted-streaming-docs-streaming-self-hosted-streaming-mdx) - [Universal 3.5 Pro Realtime on LiveKit](#universal-3-5-pro-realtime-on-livekit-docs-voice-agents-livekit-universal-3-5-pro-mdx) - [Universal 3.5 Pro Realtime on Pipecat](#universal-3-5-pro-realtime-on-pipecat-docs-voice-agents-pipecat-universal-3-5-pro-mdx) - [Streaming Migration Guide: Universal Streaming to Universal-3.5 Pro Streaming](#streaming-migration-guide-universal-streaming-to-universal-3-5-pro-streaming-docs-streaming-migration-guides-universal-to-universal-3-5-pro-streaming-mdx) - [Streaming Migration Guide: Deepgram to AssemblyAI](#streaming-migration-guide-deepgram-to-assemblyai-docs-streaming-migration-guides-dg-to-aai-streaming-mdx) - [Migration guide: Speechmatics to AssemblyAI](#migration-guide-speechmatics-to-assemblyai-docs-streaming-migration-guides-speechmatics-to-aai-streaming-mdx) - [Streaming Migration Guide: Gladia to AssemblyAI](#streaming-migration-guide-gladia-to-assemblyai-docs-streaming-migration-guides-gladia-to-aai-streaming-mdx) - [Best Practices for building Meeting Notetakers](#best-practices-for-building-meeting-notetakers-docs-meeting-notetaker-best-practices-mdx) - [Build a medical scribe](#build-a-medical-scribe-docs-medical-scribe-best-practices-mdx) - [Best Practices for Building Voice Agents](#best-practices-for-building-voice-agents-docs-voice-agents-best-practices-mdx) - [Best Practices for building Contact Center Applications](#best-practices-for-building-contact-center-applications-docs-contact-center-best-practices-mdx) - [Stream a pre-recorded file in real time](#stream-a-pre-recorded-file-in-real-time-docs-streaming-guides-stream-prerecorded-file-realtime-mdx) - [Transcribe System Audio in Real-Time (macOS)](#transcribe-system-audio-in-real-time-macos-docs-streaming-guides-transcribe-system-audio-mdx) - [Terminate Streaming Session After Inactivity](#terminate-streaming-session-after-inactivity-docs-streaming-guides-terminate-realtime-programmatically-mdx) - [Migrating from Streaming v2 to Streaming v3 (Python)](#migrating-from-streaming-v2-to-streaming-v3-python-docs-streaming-guides-v2-to-v3-migration-mdx) - [Migrating from Streaming v2 to Streaming v3 (JavaScript)](#migrating-from-streaming-v2-to-streaming-v3-javascript-docs-streaming-guides-v2-to-v3-migration-js-mdx) - [Next.js example using Real-time STT](#next-js-example-using-real-time-stt-docs-nav-links-next-js-example-using-streaming-stt-mdx) - [Vanilla Javascript example using Real-time STT](#vanilla-javascript-example-using-real-time-stt-docs-nav-links-vanilla-javascript-example-using-streaming-stt-mdx) - [Apply LLM Gateway to Streaming](#apply-llm-gateway-to-streaming-docs-streaming-guides-real-time-llm-gateway-mdx) - [Translate Real-time STT Transcripts with LLM Gateway](#translate-real-time-stt-transcripts-with-llm-gateway-docs-streaming-guides-real-time-translation-mdx) - [Apply Noise Reduction to Audio for Streaming Speech-to-Text](#apply-noise-reduction-to-audio-for-streaming-speech-to-text-docs-streaming-guides-noise-reduction-streaming-mdx) - [Transcribe audio files with Streaming](#transcribe-audio-files-with-streaming-docs-streaming-guides-streaming-transcribe-audio-file-mdx) - [Evaluate Streaming transcription accuracy with WER](#evaluate-streaming-transcription-accuracy-with-wer-docs-streaming-guides-evaluate-streaming-wer-mdx) - [Determine Optimal Turn Detection Settings from Historical Audio Analysis](#determine-optimal-turn-detection-settings-from-historical-audio-analysis-docs-streaming-guides-turn-detection-improvement-using-async-mdx) - [Transcribe a short audio file](#transcribe-a-short-audio-file-docs-sync-stt-getting-started-transcribe-a-short-audio-file-mdx) - [Prompting and Keyterms](#prompting-and-keyterms-docs-sync-stt-prompting-and-keyterms-mdx) - [Conversation context](#conversation-context-docs-sync-stt-conversation-context-mdx) - [Language selection](#language-selection-docs-sync-stt-language-selection-mdx) - [Word timestamps](#word-timestamps-docs-sync-stt-word-timestamps-mdx) - [Transcribe a short audio file](#transcribe-a-short-audio-file-docs-api-reference-sync-api-transcribe-mdx) - [Pre-warm a connection](#pre-warm-a-connection-docs-api-reference-sync-api-warm-mdx) - [Audio requirements](#audio-requirements-docs-sync-stt-audio-requirements-mdx) - [Error handling](#error-handling-docs-sync-stt-error-handling-mdx) - [Cloud endpoints & data residency](#cloud-endpoints-data-residency-docs-sync-stt-endpoints-and-data-zones-mdx) - [Connection pre-warming](#connection-pre-warming-docs-sync-stt-connection-pre-warming-mdx) - [Voice Agent API](#voice-agent-api-docs-voice-agents-voice-agent-api-mdx) - [Build with AI coding tools](#build-with-ai-coding-tools-docs-voice-agents-voice-agent-api-build-with-ai-tools-mdx) - [Supported languages](#supported-languages-docs-voice-agents-voice-agent-api-supported-languages-mdx) - [Create an agent](#create-an-agent-docs-voice-agents-voice-agent-api-create-agent-mdx) - [Manage agents](#manage-agents-docs-voice-agents-voice-agent-api-manage-agents-mdx) - [Create an agent](#create-an-agent-docs-voice-agents-voice-agent-api-api-spec-create-agent-mdx) - [List agents](#list-agents-docs-voice-agents-voice-agent-api-api-spec-list-agents-mdx) - [Retrieve an agent](#retrieve-an-agent-docs-voice-agents-voice-agent-api-api-spec-get-agent-mdx) - [Update an agent](#update-an-agent-docs-voice-agents-voice-agent-api-api-spec-update-agent-mdx) - [Delete an agent](#delete-an-agent-docs-voice-agents-voice-agent-api-api-spec-delete-agent-mdx) - [Connect your own LLM](#connect-your-own-llm-docs-voice-agents-voice-agent-api-connect-your-own-llm-mdx) - [Prompting guide](#prompting-guide-docs-voice-agents-voice-agent-api-prompting-guide-mdx) - [Greeting](#greeting-docs-voice-agents-voice-agent-api-greeting-mdx) - [Tools](#tools-docs-voice-agents-voice-agent-api-tools-overview-mdx) - [HTTP tools](#http-tools-docs-voice-agents-voice-agent-api-tools-http-tools-mdx) - [Client-side tools](#client-side-tools-docs-voice-agents-voice-agent-api-tools-client-side-tools-mdx) - [Voices](#voices-docs-voice-agents-voice-agent-api-voices-mdx) - [Output volume](#output-volume-docs-voice-agents-voice-agent-api-volume-mdx) - [Turn detection and interruptions](#turn-detection-and-interruptions-docs-voice-agents-voice-agent-api-turn-detection-and-interruptions-mdx) - [Add transcription context](#add-transcription-context-docs-voice-agents-voice-agent-api-transcription-prompt-mdx) - [Steer toward known languages](#steer-toward-known-languages-docs-voice-agents-voice-agent-api-language-selection-mdx) - [Isolate the caller's voice](#isolate-the-caller-s-voice-docs-voice-agents-voice-agent-api-noise-suppression-mdx) - [Deploy your agent](#deploy-your-agent-docs-voice-agents-voice-agent-api-deploy-mdx) - [Browser integration](#browser-integration-docs-voice-agents-voice-agent-api-browser-integration-mdx) - [Set up an inbound phone agent via SIP](#set-up-an-inbound-phone-agent-via-sip-docs-voice-agents-voice-agent-api-connect-to-twilio-mdx) - [Audio format](#audio-format-docs-voice-agents-voice-agent-api-audio-format-mdx) - [Recordings & transcripts](#recordings-transcripts-docs-voice-agents-voice-agent-api-session-history-mdx) - [List sessions](#list-sessions-docs-voice-agents-voice-agent-api-api-spec-list-sessions-mdx) - [Retrieve a session](#retrieve-a-session-docs-voice-agents-voice-agent-api-api-spec-get-session-mdx) - [Delete a session](#delete-a-session-docs-voice-agents-voice-agent-api-api-spec-delete-session-mdx) - [Inline session configuration](#inline-session-configuration-docs-voice-agents-voice-agent-api-session-configuration-mdx) - [Events reference](#events-reference-docs-voice-agents-voice-agent-api-events-reference-mdx) - [Troubleshooting](#troubleshooting-docs-voice-agents-voice-agent-api-troubleshooting-mdx) - [Message Sequence](#message-sequence-docs-voice-agents-voice-agent-api-message-sequence-mdx) - [Voice Agent WebSocket API](#voice-agent-websocket-api-docs-voice-agents-voice-agent-api-api-spec-voice-agent-websocket-mdx) - [Generate voice agent token](#generate-voice-agent-token-docs-voice-agents-voice-agent-api-api-spec-generate-voice-agent-token-mdx) - [Speech Understanding](#speech-understanding-docs-speech-understanding-getting-started-mdx) - [Action Items](#action-items-docs-speech-understanding-action-items-mdx) - [Rate limits](#rate-limits-docs-speech-understanding-rate-limits-mdx) - [Auto Chapters](#auto-chapters-docs-speech-understanding-auto-chapters-mdx) - [Custom Formatting](#custom-formatting-docs-speech-understanding-custom-formatting-mdx) - [Entity Detection](#entity-detection-docs-speech-understanding-entity-detection-mdx) - [Key Phrases](#key-phrases-docs-speech-understanding-key-phrases-mdx) - [Sentiment Analysis](#sentiment-analysis-docs-speech-understanding-sentiment-analysis-mdx) - [Speaker Identification](#speaker-identification-docs-speech-understanding-speaker-identification-mdx) - [Summarization](#summarization-docs-speech-understanding-summarization-mdx) - [Topic Detection](#topic-detection-docs-speech-understanding-topic-detection-mdx) - [Translation](#translation-docs-speech-understanding-translation-mdx) - [Migration guide: Auto Chapters](#migration-guide-auto-chapters-docs-speech-understanding-migration-guides-auto-chapters-mdx) - [Migration guide: Summarization](#migration-guide-summarization-docs-speech-understanding-migration-guides-summarization-mdx) - [Guardrails](#guardrails-docs-guardrails-getting-started-mdx) - [PII Redaction](#pii-redaction-docs-guardrails-redact-pii-from-transcripts-mdx) - [Content Moderation](#content-moderation-docs-guardrails-detect-sensitive-content-mdx) - [Profanity Filtering](#profanity-filtering-docs-guardrails-filter-profanity-from-transcripts-mdx) - [Speech Threshold](#speech-threshold-docs-guardrails-set-minimum-speech-threshold-mdx) - [Quickstart](#quickstart-docs-llm-gateway-quickstart-mdx) - [Models](#models-docs-llm-gateway-available-models-mdx) - [Providers](#providers-docs-llm-gateway-providers-mdx) - [Model Fallbacks](#model-fallbacks-docs-llm-gateway-fallback-mdx) - [Prompt Caching](#prompt-caching-docs-llm-gateway-prompt-caching-mdx) - [Tool Calling](#tool-calling-docs-llm-gateway-tool-calling-mdx) - [Structured Outputs](#structured-outputs-docs-llm-gateway-structured-outputs-mdx) - [Cloud Endpoints and Data Residency](#cloud-endpoints-and-data-residency-docs-llm-gateway-cloud-endpoints-and-data-residency-mdx) - [Create a chat completion](#create-a-chat-completion-docs-llm-gateway-api-reference-create-chat-completion-mdx) - [Create speech understanding](#create-speech-understanding-docs-llm-gateway-api-reference-create-speech-understanding-mdx) - [List available models](#list-available-models-docs-llm-gateway-api-reference-list-available-models-mdx) - [Improve Latency](#improve-latency-docs-llm-gateway-improve-latency-mdx) - [Ask Questions About Your Audio Transcripts](#ask-questions-about-your-audio-transcripts-docs-llm-gateway-ask-questions-mdx) - [Agentic Workflows](#agentic-workflows-docs-llm-gateway-agentic-workflows-mdx) - [Basic Chat Completions](#basic-chat-completions-docs-llm-gateway-chat-completions-mdx) - [Multi-turn Conversations](#multi-turn-conversations-docs-llm-gateway-conversations-mdx) - [Prompt Engineering for LLM Gateway](#prompt-engineering-for-llm-gateway-docs-llm-gateway-prompt-engineering-mdx) - [Rate limits](#rate-limits-docs-llm-gateway-rate-limits-mdx) - [Troubleshooting](#troubleshooting-docs-llm-gateway-troubleshooting-mdx) - [Apply LLM Gateway to Audio Transcripts](#apply-llm-gateway-to-audio-transcripts-docs-llm-gateway-apply-llms-to-audio-files-mdx) - [Setup An AI Coach With LLM Gateway](#setup-an-ai-coach-with-llm-gateway-docs-guides-task-endpoint-ai-coach-mdx) - [Generate Action Items with LLM Gateway](#generate-action-items-with-llm-gateway-docs-guides-task-endpoint-action-items-mdx) - [Prompt A Structured Q&A Response Using LLM Gateway](#prompt-a-structured-q-a-response-using-llm-gateway-docs-guides-task-endpoint-structured-qa-mdx) - [Estimate Input Token Costs for LLM Gateway](#estimate-input-token-costs-for-llm-gateway-docs-guides-counting-tokens-mdx) - [Extract Dialogue Data with LLM Gateway and JSON](#extract-dialogue-data-with-llm-gateway-and-json-docs-guides-dialogue-data-mdx) - [Extract Quotes with Timestamps Using LLM Gateway + Semantic Search](#extract-quotes-with-timestamps-using-llm-gateway-semantic-search-docs-guides-transcript-citations-mdx) - [Extract Transcript Quotes with LLM Gateway](#extract-transcript-quotes-with-llm-gateway-docs-guides-timestamped-transcripts-mdx) - [Analyze The Sentiment Of A Customer Call using LLM Gateway](#analyze-the-sentiment-of-a-customer-call-using-llm-gateway-docs-guides-call-sentiment-analysis-mdx) - [Custom Topic Tags Using LLM Gateway](#custom-topic-tags-using-llm-gateway-docs-guides-custom-topic-tags-mdx) - [Redact PII from Text Using LLM Gateway](#redact-pii-from-text-using-llm-gateway-docs-guides-llm-gateway-pii-redaction-mdx) - [Apply LLM Gateway to Streaming](#apply-llm-gateway-to-streaming-docs-guides-real-time-llm-gateway-mdx) - [Translate Real-time STT Transcripts with LLM Gateway](#translate-real-time-stt-transcripts-with-llm-gateway-docs-guides-real-time-translation-mdx) - [Implement a Sales Playbook Using LLM Gateway](#implement-a-sales-playbook-using-llm-gateway-docs-guides-sales-playbook-mdx) - [Segment A Phone Call using LLM Gateway](#segment-a-phone-call-using-llm-gateway-docs-guides-phone-call-segmentation-mdx) - [Generate SOAP Notes using LLM Gateway](#generate-soap-notes-using-llm-gateway-docs-guides-soap-note-generation-mdx) - [Frequently Asked Questions](#frequently-asked-questions-docs-faq-mdx) - [Has AssemblyAI certified to the EU-U.S. Data Privacy Framework?](#has-assemblyai-certified-to-the-eu-u-s-data-privacy-framework-docs-faq-has-assemblyai-certified-to-the-eu-us-data-privacy-framework-mdx) - [Can I sign a DPA agreement with AssemblyAI?](#can-i-sign-a-dpa-agreement-with-assemblyai-docs-faq-can-i-sign-a-dpa-agreement-with-assemblyai-mdx) - [Can you provide a copy of your most recent penetration test executive summary?](#can-you-provide-a-copy-of-your-most-recent-penetration-test-executive-summary-docs-faq-can-you-provide-a-copy-of-your-most-recent-penetration-test-executive-summary-mdx) - [Can you provide a recent vulnerability scan?](#can-you-provide-a-recent-vulnerability-scan-docs-faq-can-you-provide-a-recent-vulnerability-scan-mdx) - [Will AssemblyAI sign a Business Associate Addendum (BAA) as described in the HIPAA rules and regulations?](#will-assemblyai-sign-a-business-associate-addendum-baa-as-described-in-the-hipaa-rules-and-regulations-docs-faq-can-you-sign-a-baa-mdx) - [Do you have a formal risk assessment policy or process?](#do-you-have-a-formal-risk-assessment-policy-or-process-docs-faq-do-you-have-a-formal-risk-assessment-policy-or-process-mdx) - [Do you have documented information security policies? If so, how frequently are they updated?](#do-you-have-documented-information-security-policies-if-so-how-frequently-are-they-updated-docs-faq-do-you-have-documented-information-security-policies-if-so-how-frequently-are-they-updated-mdx) - [Do you offer EU Data Residency?](#do-you-offer-eu-data-residency-docs-faq-do-you-offer-eu-data-residency-mdx) - [Do you offer self-hosted solutions?](#do-you-offer-self-hosted-solutions-docs-faq-do-you-offer-self-hosted-solutions-mdx) - [Do you offer servers in the EU?](#do-you-offer-servers-in-the-eu-docs-faq-do-you-offer-servers-in-the-eu-mdx) - [Do you support SAML in your product?](#do-you-support-saml-in-your-product-docs-faq-do-you-support-saml-in-your-product-mdx) - [How long are outputs maintained?](#how-long-are-outputs-maintained-docs-faq-how-long-are-outputs-retained-mdx) - [Does AssemblyAI utilize an anti-virus/anti-malware solution across all relevant infrastructure (workstations and servers), and are appropriate response capabilities deployed to respond to alerts?](#does-assemblyai-utilize-an-anti-virus-anti-malware-solution-across-all-relevant-infrastructure-workstations-and-servers-and-are-appropriate-response-capabilities-deployed-to-respond-to-alerts-docs-faq-does-assemblyai-utilize-an-anti-virus-anti-malware-solution-across-all-relevant-infrastructure-workstations-and-servers-and-are-appropriate-response-capabilities-deployed-to-respond-to-ale-mdx) - [How are incidents escalated within your organization?](#how-are-incidents-escalated-within-your-organization-docs-faq-how-are-incidents-escalated-within-your-organization-mdx) - [How do we securely use your service?](#how-do-we-securely-use-your-service-docs-faq-how-do-we-securely-use-your-service-mdx) - [How do you protect production code?](#how-do-you-protect-production-code-docs-faq-how-do-you-protect-production-code-mdx) - [How to Access AssemblyAI's Security Reports](#how-to-access-assemblyai-s-security-reports-docs-faq-how-to-access-assemblyai-s-security-reports-mdx) - [How to Opt Out of Data Sharing for our Model Improvement Program](#how-to-opt-out-of-data-sharing-for-our-model-improvement-program-docs-faq-how-to-opt-out-of-data-sharing-for-our-model-improvement-program-mdx) - [Is multi-factor authentication enforced for all access to scoped systems and data?](#is-multi-factor-authentication-enforced-for-all-access-to-scoped-systems-and-data-docs-faq-is-multi-factor-authentication-enforced-for-all-access-to-scoped-systems-and-data-mdx) - [Does AssemblyAI have a documented process for reviewing and approving third-party service providers?](#does-assemblyai-have-a-documented-process-for-reviewing-and-approving-third-party-service-providers-docs-faq-is-there-a-documented-process-for-reviewing-and-approving-third-party-service-providers-mdx) - [Does AssemblyAI have an incident response plan?](#does-assemblyai-have-an-incident-response-plan-docs-faq-please-describe-the-incident-response-plan-mdx) - [What are your recovery time and recovery point objectives?](#what-are-your-recovery-time-and-recovery-point-objectives-docs-faq-what-are-your-recovery-time-and-recovery-point-objectives-mdx) - [What is your SLA for repairing Critical/High/Medium vulnerabilities?](#what-is-your-sla-for-repairing-critical-high-medium-vulnerabilities-docs-faq-what-is-your-sla-for-repairing-critical-high-medium-vulnerabilities-mdx) - [What logs are available to customers?](#what-logs-are-available-to-customers-docs-faq-what-logs-are-available-to-customers-mdx) - [What standards do your internal password policies follow?](#what-standards-do-your-internal-password-policies-follow-docs-faq-what-standards-do-your-internal-password-policies-follow-mdx) - [Where are your servers located?](#where-are-your-servers-located-docs-faq-where-are-your-servers-located-mdx) - [Am I charged for transcribing silent audio?](#am-i-charged-for-transcribing-silent-audio-docs-faq-am-i-charged-for-transcribing-silent-audio-mdx) - [Are Custom Models More Accurate than General Models?](#are-custom-models-more-accurate-than-general-models-docs-faq-are-custom-models-more-accurate-than-general-models-mdx) - [Do I Get Charged for Failed API Calls?](#do-i-get-charged-for-failed-api-calls-docs-faq-are-customers-charged-for-api-calls-that-result-in-errors-mdx) - [Are there any limits on file size or file duration for files submitted to the API?](#are-there-any-limits-on-file-size-or-file-duration-for-files-submitted-to-the-api-docs-faq-are-there-any-limits-on-file-size-or-file-duration-for-files-submitted-to-the-api-mdx) - [Can I customize how words are spelled by the model?](#can-i-customize-how-words-are-spelled-by-the-model-docs-faq-can-i-customize-how-words-are-spelled-by-the-model-mdx) - [Can I delete the transcripts I have created using the API?](#can-i-delete-the-transcripts-i-have-created-using-the-api-docs-faq-can-i-delete-the-transcripts-i-have-created-using-the-api-mdx) - [Can I get a list of all transcripts I have created?](#can-i-get-a-list-of-all-transcripts-i-have-created-docs-faq-can-i-get-a-list-of-all-transcripts-i-have-created-mdx) - [Can I send audio to AssemblyAI in segments and still get speaker labels for the whole recording?](#can-i-send-audio-to-assemblyai-in-segments-and-still-get-speaker-labels-for-the-whole-recording-docs-faq-can-i-send-audio-to-assemblyai-in-segments-and-still-get-speaker-labels-for-the-whole-recording-mdx) - [Can I submit files to the API that are stored in a Google Drive?](#can-i-submit-files-to-the-api-that-are-stored-in-a-google-drive-docs-faq-can-i-submit-files-to-the-api-that-are-stored-in-a-google-drive-mdx) - [Can I use the API without internet access?](#can-i-use-the-api-without-internet-access-docs-faq-can-i-use-the-api-without-internet-access-mdx) - [Do we have resources for building with Make?](#do-we-have-resources-for-building-with-make-docs-faq-do-we-have-resources-for-building-with-make-mdx) - [Do you have any examples for how to use your API?](#do-you-have-any-examples-for-how-to-use-your-api-docs-faq-do-you-have-any-examples-for-how-to-use-your-api-mdx) - [Do you have example use cases for using AssemblyAI?](#do-you-have-example-use-cases-for-using-assemblyai-docs-faq-do-you-have-example-use-cases-for-using-assemblyai-mdx) - [Do you offer cross-file Speaker Identification?](#do-you-offer-cross-file-speaker-identification-docs-faq-do-you-offer-cross-file-speaker-identification-mdx) - [Do you offer translation?](#do-you-offer-translation-docs-faq-do-you-offer-translation-mdx) - [Do you offer voice-to-voice or text-to-speech (TTS)?](#do-you-offer-voice-to-voice-or-text-to-speech-tts-docs-faq-do-you-offer-voice-to-voice-or-text-to-speech-tts-mdx) - [Does it cost extra to export SRT or VTT captions?](#does-it-cost-extra-to-export-srt-or-vtt-captions-docs-faq-does-it-cost-extra-to-export-srt-or-vtt-captions-mdx) - [Is there a way to generate SRT or VTT captions with speaker labels?](#is-there-a-way-to-generate-srt-or-vtt-captions-with-speaker-labels-docs-faq-is-there-a-way-to-generate-srt-or-vtt-captions-with-speaker-labels-mdx) - [Does it cost more to transcribe an audio or video?](#does-it-cost-more-to-transcribe-an-audio-or-video-docs-faq-does-it-cost-more-to-transcribe-an-audio-or-video-mdx) - [Does your API return timestamps for individual words?](#does-your-api-return-timestamps-for-individual-words-docs-faq-does-your-api-return-timestamps-for-individual-words-mdx) - [How are individual speakers identified and how does the Speaker Label feature work?](#how-are-individual-speakers-identified-and-how-does-the-speaker-label-feature-work-docs-faq-how-are-individual-speakers-identified-and-how-does-the-speaker-label-feature-work-mdx) - [How are paragraphs created for the /paragraphs endpoint?](#how-are-paragraphs-created-for-the-paragraphs-endpoint-docs-faq-how-are-paragraphs-created-for-the-paragraphs-endpoint-mdx) - [How are word/transcript level confidence scores calculated?](#how-are-word-transcript-level-confidence-scores-calculated-docs-faq-how-are-word-transcript-level-confidence-scores-calculated-mdx) - [How can I integrate AssemblyAI with other services?](#how-can-i-integrate-assemblyai-with-other-services-docs-faq-how-can-i-integrate-assemblyai-with-other-services-mdx) - [How can I make certain words more likely to be transcribed?](#how-can-i-make-certain-words-more-likely-to-be-transcribed-docs-faq-how-can-i-make-certain-words-more-likely-to-be-transcribed-mdx) - [How can I test AssemblyAI without writing code?](#how-can-i-test-assemblyai-without-writing-code-docs-faq-how-can-i-test-assemblyai-without-writing-code-mdx) - [How can I transcribe YouTube videos?](#how-can-i-transcribe-youtube-videos-docs-faq-how-can-i-transcribe-youtube-videos-mdx) - [How do I generate subtitles?](#how-do-i-generate-subtitles-docs-faq-how-do-i-generate-subtitles-mdx) - [How does AssemblyAI compare to other ASR providers?](#how-does-assemblyai-compare-to-other-asr-providers-docs-faq-how-does-assemblyai-compare-to-other-asr-providers-mdx) - [How does Automatic Language Detection work?](#how-does-automatic-language-detection-work-docs-faq-how-does-language-detection-work-for-transcriptions-mdx) - [How does the API handle files that contain spoken audio in multiple languages?](#how-does-the-api-handle-files-that-contain-spoken-audio-in-multiple-languages-docs-faq-how-does-the-api-handle-files-that-contain-spoken-audio-in-multiple-languages-mdx) - [How long does it take to transcribe a file?](#how-long-does-it-take-to-transcribe-a-file-docs-faq-how-long-does-it-take-to-transcribe-a-file-mdx) - [What should I do if I'm getting an error?](#what-should-i-do-if-i-m-getting-an-error-docs-faq-i-am-getting-an-error-what-should-i-do-mdx) - [Is there a Postman collection for using the API?](#is-there-a-postman-collection-for-using-the-api-docs-faq-is-there-a-postman-collection-for-using-the-api-mdx) - [Is there a way for us to send the start time / end time for transcription instead of transcribing the whole length of a call recording?](#is-there-a-way-for-us-to-send-the-start-time-end-time-for-transcription-instead-of-transcribing-the-whole-length-of-a-call-recording-docs-faq-is-there-a-way-for-us-to-send-the-start-time-end-time-for-transcription-instead-of-transcribing-the-whole-length-of-a-call-recording-mdx) - [Is there an OpenAPI spec/schema for the API?](#is-there-an-openapi-spec-schema-for-the-api-docs-faq-is-there-an-openapi-spec-schema-for-the-api-mdx) - [What causes a "read operation timed out" error?](#what-causes-a-read-operation-timed-out-error-docs-faq-read-operation-timed-out-error-mdx) - [Should I use Speaker Labels or Multi-channel?](#should-i-use-speaker-labels-or-multi-channel-docs-faq-should-i-use-speaker-labels-or-multi-channel-mdx) - [Should I use the NA or EU endpoint for my Speech-to-Text requests?](#should-i-use-the-na-or-eu-endpoint-for-my-speech-to-text-requests-docs-faq-should-i-use-the-na-or-eu-endpoint-mdx) - [What are the recommended options for audio noise reduction?](#what-are-the-recommended-options-for-audio-noise-reduction-docs-faq-what-are-the-recommended-options-for-audio-noise-reduction-mdx) - [What audio and video file types are supported by your API?](#what-audio-and-video-file-types-are-supported-by-your-api-docs-faq-what-audio-and-video-file-types-are-supported-by-your-api-mdx) - [What IP Address Should I Whitelist for AssemblyAI?](#what-ip-address-should-i-whitelist-for-assemblyai-docs-faq-what-ip-address-should-i-whitelist-for-assemblyai-mdx) - [What is the minimum audio duration that the API can transcribe?](#what-is-the-minimum-audio-duration-that-the-api-can-transcribe-docs-faq-what-is-the-minimum-audio-duration-that-the-api-can-transcribe-mdx) - [What is the recommended file type for using your API?](#what-is-the-recommended-file-type-for-using-your-api-docs-faq-what-is-the-recommended-file-type-for-using-your-api-mdx) - [What types of audio URLs can I use with the API?](#what-types-of-audio-urls-can-i-use-with-the-api-docs-faq-what-types-of-audio-urls-can-i-use-with-the-api-mdx) - [Where can I find a list of recent changes to the API?](#where-can-i-find-a-list-of-recent-changes-to-the-api-docs-faq-where-can-i-find-a-list-of-recent-changes-to-the-api-mdx) - [Where can I find cURL code examples?](#where-can-i-find-curl-code-examples-docs-faq-where-can-i-find-curl-code-examples-mdx) - [Why can't I access recording URLs from the /upload endpoint directly?](#why-can-t-i-access-recording-urls-from-the-upload-endpoint-directly-docs-faq-why-cant-i-access-urls-from-the-upload-endpoint-directly-mdx) - [Can I use speaker diarization with Streaming Speech-to-Text?](#can-i-use-speaker-diarization-with-streaming-speech-to-text-docs-faq-can-i-use-speaker-diarization-with-live-audio-transcription-mdx) - [How accurate is your Streaming transcription compared to Async transcription?](#how-accurate-is-your-streaming-transcription-compared-to-async-transcription-docs-faq-how-accurate-is-your-real-time-transcription-compared-to-async-transcription-mdx) - [How does Universal Streaming session-based pricing work?](#how-does-universal-streaming-session-based-pricing-work-docs-faq-how-does-universal-streaming-session-based-pricing-work-mdx) - [What languages are supported for Streaming Speech-to-text?](#what-languages-are-supported-for-streaming-speech-to-text-docs-faq-language-support-for-real-time-transcription-mdx) - [Resolving SSL Certificate Verification Error When Trying to Use Real-time STT](#resolving-ssl-certificate-verification-error-when-trying-to-use-real-time-stt-docs-faq-resolving-ssl-certificate-verification-error-in-assemblyai-real-time-transcriber-mdx) - [I am getting a "Model deprecated. See docs for new model information" error message. What does it mean?](#i-am-getting-a-model-deprecated-see-docs-for-new-model-information-error-message-what-does-it-mean-docs-faq-upgrading-to-the-universal-streaming-model-mdx) - [How do Content Moderation severity scores work?](#how-do-content-moderation-severity-scores-work-docs-faq-how-do-content-moderation-severity-scores-work-mdx) - [How can I summarize my audio file?](#how-can-i-summarize-my-audio-file-docs-faq-how-do-your-summarization-models-work-mdx) - [Is Mistral still supported?](#is-mistral-still-supported-docs-faq-is-mistral-still-supported-mdx) - [Is pricing for Speech Understanding per feature or all-inclusive?](#is-pricing-for-speech-understanding-per-feature-or-all-inclusive-docs-faq-is-pricing-for-audio-intelligence-per-feature-or-all-inclusive-mdx) - [Understanding Input and Output Tokens for LLM Gateway](#understanding-input-and-output-tokens-for-llm-gateway-docs-faq-understanding-input-and-output-tokens-for-llm-gateway-mdx) - [Can you use the Playground with files in languages other than English?](#can-you-use-the-playground-with-files-in-languages-other-than-english-docs-faq-can-you-use-the-playground-with-files-in-languages-other-than-english-mdx) - [How do I delete a transcript I created using the Playground?](#how-do-i-delete-a-transcript-i-created-using-the-playground-docs-faq-how-do-i-delete-a-transcript-i-created-using-the-playground-mdx) - [Why is the transcription I am receiving using the Playground in a different language?](#why-is-the-transcription-i-am-receiving-using-the-playground-in-a-different-language-docs-faq-why-is-the-transcription-i-am-receiving-using-the-playground-in-a-different-language-mdx) - [Do you have an affiliate marketing program?](#do-you-have-an-affiliate-marketing-program-docs-faq-do-you-have-an-affiliate-marketing-program-mdx) - [Do you have any job openings or internship opportunities?](#do-you-have-any-job-openings-or-internship-opportunities-docs-faq-do-you-have-any-job-openings-or-internship-opportunities-mdx) - [How do I contact support?](#how-do-i-contact-support-docs-faq-how-do-i-contact-support-mdx) - [How do I get in touch with your Sales team?](#how-do-i-get-in-touch-with-your-sales-team-docs-faq-how-do-i-get-in-touch-with-your-sales-team-mdx) - [I’ve spotted an issue with the website, what should I do?](#i-ve-spotted-an-issue-with-the-website-what-should-i-do-docs-faq-ive-spotted-an-issue-with-the-website-what-should-i-do-mdx) - [What are your support hours and response time SLAs?](#what-are-your-support-hours-and-response-time-slas-docs-faq-what-are-your-support-hours-and-response-time-slas-mdx) - [What is your API Uptime SLA?](#what-is-your-api-uptime-sla-docs-faq-what-is-your-api-uptime-sla-mdx) - [Where can I find AssemblyAI's product roadmap?](#where-can-i-find-assemblyai-s-product-roadmap-docs-faq-where-can-i-find-assemblyais-product-roadmap-mdx) - [AssemblyAI integration prompt for AI coding agents](#assemblyai-integration-prompt-for-ai-coding-agents-docs-agent-instructions-mdx) - [Deployment](#deployment-docs-deployment-mdx) - [Select The EU Region for EU Data Residency](#select-the-eu-region-for-eu-data-residency-docs-pre-recorded-audio-guides-how-to-use-the-eu-endpoint-mdx) - [Identifying speakers in audio recordings](#identifying-speakers-in-audio-recordings-docs-pre-recorded-audio-guides-identifying-speakers-in-audio-recordings-mdx) - [Streaming v2 (Legacy)](#streaming-v2-legacy-docs-pre-recorded-audio-legacy-streaming-mdx) - [Universal-3.5 Pro](#universal-3-5-pro-docs-pre-recorded-audio-universal-3-5-pro-mdx) - [Tracking customer transcription usage](#tracking-customer-transcription-usage-docs-tracking-your-customers-usage-mdx) - [Use case guides](#use-case-guides-docs-use-cases-mdx) - [Universal-3.5 Pro Streaming API](#universal-3-5-pro-streaming-api-docs-voice-agents-universal-3-5-pro-streaming-api-mdx) - [Create a webhook subscription](#create-a-webhook-subscription-docs-voice-agents-voice-agent-api-api-spec-create-webhook-subscription-mdx) - [Delete a webhook subscription](#delete-a-webhook-subscription-docs-voice-agents-voice-agent-api-api-spec-delete-webhook-subscription-mdx) - [Retrieve a webhook subscription](#retrieve-a-webhook-subscription-docs-voice-agents-voice-agent-api-api-spec-get-webhook-subscription-mdx) - [List webhook subscriptions](#list-webhook-subscriptions-docs-voice-agents-voice-agent-api-api-spec-list-webhook-subscriptions-mdx) - [Update a webhook subscription](#update-a-webhook-subscription-docs-voice-agents-voice-agent-api-api-spec-update-webhook-subscription-mdx) - [Webhooks](#webhooks-docs-voice-agents-voice-agent-api-webhooks-mdx) - [API specs](#api-specs) --- # AssemblyAI Documentation URL: https://www.assemblyai.com/docs Source: docs/index.mdx Navigation: Overview > Getting started Description: Build with our Voice AI Infrastructure ## Build with AI coding agents Pin the integration prompt URL in your agent's `AGENTS.md` / `CLAUDE.md` / `.cursorrules`, or copy the full prompt directly. Works with Cursor, Claude Code, Copilot, Devin, and any other AI coding assistant. ## Explore our models ## Deployment options } href="/voice-agents/livekit-universal-3-5-pro" /> } href="/voice-agents/pipecat-universal-3-5-pro" /> --- # Build with AI coding agents URL: https://www.assemblyai.com/docs/coding-agent-prompts Source: docs/coding-agent-prompts.mdx Navigation: Overview > Getting started Description: Build with AssemblyAI using Cursor, Claude Code, Copilot, Devin, and other AI coding assistants. Copy the integration prompt, pin the URL in your agent's project instructions, connect the docs MCP server, or install the Claude Code skill. If you're using Cursor, Claude Code, Copilot, ChatGPT Codex, Devin, or another AI coding assistant to build with AssemblyAI, give your agent live context so it doesn't rely on outdated training data. ## Start here Paste our full integration prompt into your agent. Teaches it operating rules, recommended models, parameter gotchas, and quickstarts for every product surface — so the first code you get back is closer to production-ready. Add a one-liner to your `AGENTS.md`, `CLAUDE.md`, or `.cursorrules` so your agent refetches the latest integration prompt before every AssemblyAI prompt — **the most effective option** and what we recommend for ongoing work. ## Compare all four methods | Method | Who it's for | When to use | Trade-off | |---|---|---|---| | [**Pin the URL in your agent instructions**](#pin-the-url-in-your-agent-instructions) (`CLAUDE.md`, `.cursorrules`, `AGENTS.md`, …) | Teams shipping AssemblyAI in production — **recommended default.** | Ongoing project work: agent reads project instructions on every prompt and refetches the latest docs. | Tiny one-liner; agent re-fetches the latest docs before every AssemblyAI prompt. | | [**Copy the integration prompt**](#copy-the-integration-prompt) | First-time users and quick experiments. | One-shot chat sessions (ChatGPT Codex, single Cursor / Claude Code conversation) where you don't have a project config to pin a URL in. | Adds ~770 lines to your prompt; re-copy when the prompt changes. Less efficient than the URL pin for ongoing work. | | [**Docs MCP server**](#docs-mcp-server) | Devs whose agent does heavy mid-session doc lookups. | Iterative feature work where the agent needs to check several pages mid-task. | Layers on top of the URL pin; requires an MCP-capable client. | | [**Claude Code skill**](#claude-code-skill) | Claude Code and other skills-compatible agents (60+, via the universal skills format). | You want curated SDK-specific context (Python / JS / streaming / voice agents) bundled with the agent, auto-loaded when relevant. | Installed snapshot; re-install to pick up doc changes. | For most teams, pinning the URL is enough. Add the MCP server on top if your agent does a lot of mid-session lookups. The copy-paste prompt is best kept for first-time use or one-shot chat sessions. ## Pin the URL in your agent instructions For ongoing projects, point your agent at the integration prompt URL from its project instructions. The agent re-fetches the prompt before every AssemblyAI-related prompt, so you automatically pick up API changes without bloating your context window. Add this to your project's `AGENTS.md`, `CLAUDE.md`, `.cursorrules`, `.cursor/rules/`, `.github/copilot-instructions.md`, `.windsurfrules`, or whichever instructions file your agent reads: ``` Before writing AssemblyAI code, read https://www.assemblyai.com/docs/agent-instructions.md and https://www.assemblyai.com/docs/llms.txt. The API has changed — do not rely on memorized parameter names. ``` The two URLs cover different things: - [**agent-instructions.md**](/agent-instructions.md) — the **rules of the road**: operating rules, discovery questions, recommendation template, feature selection, quickstarts for every product surface, error handling, and the gotchas that trip up most integrations. - [**llms.txt**](/llms.txt) — the **documentation index**: every page's title, URL, and one-line description so the agent can navigate to detail pages. For deep lookups, [llms-full.txt](/llms-full.txt) returns the full concatenated documentation content. Narrow with `?lang=python` or `?lang=typescript` to save tokens, or add `?excludeSpec=true` to skip the API reference. ## Copy the integration prompt For a one-shot agent session — a single ChatGPT Codex run, a quick Cursor chat without project files, or just trying AssemblyAI for the first time — paste the full integration prompt into your agent along with a short description of what you want to build. The prompt teaches the agent the operating rules, discovery questions, recommended models, parameter gotchas, and quickstarts for every product surface, so the first code you get back is closer to production-ready than what the agent would write from memory. For ongoing project work, [pin the URL](#pin-the-url-in-your-agent-instructions) instead — pasting the full prompt into your project instructions stuffs every message with ~770 lines of context. ````` # AssemblyAI Integration — Coding Agent Instructions You are helping a developer integrate AssemblyAI's Speech-to-Text API into their application. Your job is to understand their context through discovery, produce a concrete implementation plan, get their approval, and then write correct, production-ready code. This is a public API. The developer creates their own key at [assemblyai.com/dashboard/api-keys](https://www.assemblyai.com/dashboard/home). **Official documentation.** Two ways to wire your coding agent up to live docs (both recommended — they layer): 1. **Project instructions** (every prompt): add to `CLAUDE.md`, `.cursorrules`, `AGENTS.md`, or equivalent: ``` Always fetch https://www.assemblyai.com/docs/llms.txt before writing AssemblyAI code. The API has changed — do not rely on memorized parameter names. ``` `llms.txt` is the structured index. For full content use `llms-full.txt`; narrow with `?lang=python` or `?lang=typescript`, or add `?excludeSpec=true` to skip the API spec. 2. **Docs MCP server** (on-demand lookups): `https://assemblyai.com/docs/mcp` — Streamable HTTP transport. Lets the agent search and fetch AssemblyAI documentation pages on demand. ```bash # Claude Code claude mcp add assemblyai-docs --transport http https://assemblyai.com/docs/mcp ``` See the [Coding agent prompts](/coding-agent-prompts) page for Cursor and other clients. --- ## 0. Operating rules 1. **Discovery first, code later.** Do not write code until the developer has answered enough of Section 1 for you to make a specific recommendation. 2. **One question per message.** Never batch discovery questions. Wait for an answer before asking the next one. 3. **Plan before you build.** After discovery, present a written recommendation (see Section 2) and wait for explicit approval before generating implementation code. 4. **Prefer the official SDKs.** Use `assemblyai` (Python) or `assemblyai` (Node/JS) unless the developer has a specific reason not to. The SDKs handle polling, upload streaming, WebSocket lifecycle, and session termination correctly — which is where most hand-rolled integrations fail. 5. **Never expose the API key in client-side code.** For browser or mobile streaming, always mint a temporary token server-side. For pre-recorded, proxy uploads and submissions through your server. 6. **Authorization header is the raw key — no `Bearer` prefix.** This trips up everyone. **One exception:** the Voice Agent API (Section 10) requires `Authorization: Bearer YOUR_API_KEY`. Don't generalize either rule across products. 7. **`speech_models` is optional on pre-recorded requests.** If omitted, the request defaults to `["universal-3-5-pro", "universal-2"]`. You can still set it explicitly — see Section 5 for semantics. 8. **Always terminate streaming sessions explicitly.** An abandoned WebSocket keeps accruing charges until the 3-hour cap. 9. **Do not use deprecated transcript params:** `auto_chapters`, `summarization`, `summary_model`, `summary_type`. Use LLM Gateway instead (Section 8). 10. **If the developer's answers are inconsistent, stop and surface the conflict.** Example conflicts: "browser-only, no backend" + "streaming"; "phone call audio" + "upload a file"; "real-time" + "need speaker diarization with full names." Don't paper over these — ask. 11. **Be flexible.** If something the developer says doesn't match the shape of the API (e.g., they describe a use case that isn't supported — see Section 13), say so directly and propose the closest supported alternative. 12. **Verify parameters against live docs before recommending.** This file is a snapshot — features move between beta and GA, model-specific behaviors change, and new knobs ship regularly. Before posting the Section 2 recommendation, confirm each parameter you plan to use is supported for the chosen **mode** (pre-recorded vs streaming) *and* **model** (U3 Pro, U2, U3 Pro Streaming, Universal-Streaming). Do not assume a pre-recorded flag works on streaming, or that a parameter supported on U2 still behaves the same on U3 Pro. Pull the current reference rather than memorizing. Primary sources, in order of preference: - `https://www.assemblyai.com/docs/llms-full.txt` — the canonical machine-readable reference - Per-mode docs: `/docs/pre-recorded-audio/*` (pre-recorded) and `/docs/streaming/*` (streaming), including the model-specific overview page (e.g., `/docs/streaming/getting-started/transcribe-streaming-audio` and `/docs/streaming/select-the-speech-model`) which lists *exactly* which parameters are honored/ignored by that model - The OpenAPI-backed API reference at `/docs/api-reference/*` for request/response schemas - For LLM Gateway: `/docs/llm-gateway/quickstart` lists the current valid `model` strings — don't guess short names like `claude-sonnet-4` If a flag you remembered isn't in the current docs (or is marked beta / deprecated / ignored for the chosen model), flag it in the recommendation's "Open questions / assumptions" block and ask the developer before proceeding. --- ## 1. Discovery questions Ask these **one at a time**, in order. Skip any question already answered in the conversation. Adapt wording to sound natural, but cover the substance of each. 1. **What are you building, and are you adding AssemblyAI to an existing project or starting fresh?** (A short description of the product is usually enough.) 2. **What do you need: pre-recorded transcription, real-time real-time STT, or a managed voice agent?** - Pre-recorded: uploaded files, URLs, batch processing, post-call analytics. → Section 6. - Real-time STT: live transcripts only (you bring your own LLM/TTS). Live captioning, voice-agent STT, meeting notetaking, dictation. → Section 9. - Voice Agent API (managed): full-duplex speech-in/speech-out — STT + LLM + TTS + turn detection + tool calling, all in one WebSocket. Right answer when "I want to talk to an AI" is the whole product. → Section 10. 3. **Where is your audio coming from?** (e.g., uploaded files, public URLs, browser microphone, mobile app, Twilio/Telnyx phone numbers, SIP trunks.) 4. **What language and framework are you using?** (e.g., Python + FastAPI, Node + Next.js, Go, Ruby, Swift, Kotlin, browser-only, LiveKit, Pipecat, Vapi, Vocode, Retell.) 5. **Do you already have an AssemblyAI API key, or do you need to create one?** (If needed: [assemblyai.com/dashboard/api-keys](https://www.assemblyai.com/dashboard/home).) 6. **Do you have a data residency requirement?** (US vs EU — this changes the base URL.) 7. **Anything beyond a plain transcript?** Don't read off a checklist. Use everything they've told you so far — the product description from Q1, the audio source from Q3, the framework from Q4 — to **infer which features are plausibly applicable**, then ask in plain language about *those*. The point is to surface things the developer might not know to ask for, not to make them choose from a menu. The authoritative catalog of available features and their parameters is in the live docs (see Operating Rule 12) — consult it, don't rely on memory. Section 3 of this file is a starting reference, not the final word. Calibrate to mode and use case. Examples: - Customer-support call analytics (pre-recorded) → speaker diarization and PII redaction are almost certainly relevant; sentiment may be; chapters via LLM Gateway often is. Ask about those, not about live-streaming features. - Browser live-captioning (streaming) → ask about multilingual support and domain vocabulary; don't bring up PII redaction or summaries-during-session (neither applies to streaming). - Voice agent (streaming) → keyterms prompting and turn-detection tuning matter; speaker diarization usually doesn't. - Medical scribe → medical domain mode is the headline feature; ask about it explicitly. Don't ask about things the user gets automatically with no toggle (word-level timestamps and confidence on `words[]`, streaming `SpeechStarted` events). Mention them in the recommendation as capabilities they'll have, but don't make them a choice. If you're confident from context that a feature is needed (e.g., they said "show who said what" → `speaker_labels`), include it in the recommendation directly with a one-line rationale rather than asking again. --- ## 2. Recommendation template (after discovery) Before writing code, post a plan with all of the following. Get explicit approval. ```` ## Recommendation **Use case:** **Mode:** **Region:** **Model:** - - **Endpoints:** - - **Parameters enabled:** (before filling this in, verify each parameter is supported on the chosen mode + model per Operating Rule 12) - `param_name`: - ... **Auth pattern:** **Termination & error handling:** **Code skeleton:** <2–6 bullet points describing the files/functions you'll generate> **Open questions / assumptions:** Ready to proceed? ```` If they say yes, write the code. If they push back on any piece, revise the plan — don't just start coding around objections. --- ## 3. Feature selection guide (agent reference) Use this to build the recommendation. Do not dump it on the user. | Developer need | Parameter / approach | |---|---| | Speaker diarization | `speaker_labels: true` (pre-recorded, and streaming — streaming adds a `speaker_label` to each Turn event) | | Automatic language detection | `language_detection: true` (pre-recorded; on streaming, only available on Universal-Streaming Multilingual — adds `language_code` + `language_confidence` to Turn events. **Not** supported on U3 Pro Streaming.) | | Specific language | `language_code: "es"` etc. (pre-recorded and U3 Pro Streaming — biases the multilingual model toward one language) | | Multilingual / code-switching | `speech_models: ["universal-3-5-pro"]` + `prompt` parameter — see [U3 Pro prompting guide](/pre-recorded-audio/universal-3-5-pro/prompting) | | Domain-specific vocabulary | `keyterms_prompt: [...]` (pre-recorded: up to 1,000 terms with U3 Pro / 200 with U2; streaming: up to 100 terms, each ≤50 chars) | | Medical domain | `domain: "medical-v1"` (pre-recorded *and* streaming; supported languages: en, es, de, fr) | | PII redaction in text | `redact_pii: true` + `redact_pii_policies: [...]` + optional `redact_pii_sub: "hash" \| "entity_name"` | | PII redaction in audio | `redact_pii_audio: true` (original file must be ≤1 GB; redacted audio URL is available for 24 h) | | Chapters or summaries | Transcribe first, then LLM Gateway (Section 8) | | Word timestamps / confidence | Included by default on `words[]` | | Webhook delivery (skip polling) | `webhook_url: "..."` (Section 7) | | Managed voice agent (speech-in / speech-out) | Voice Agent API (Section 10) — one WebSocket, no separate STT/LLM/TTS | | Custom voice agent (your LLM + TTS) | Real-time STT + framework integration (Section 11) | | Multilingual streaming | Universal-3 Pro Streaming (native code-switching; `language_code` to bias toward one language) | --- ## 4. API overview - **REST base URL (US):** `https://api.assemblyai.com` - **REST base URL (EU):** `https://api.eu.assemblyai.com` - **Streaming WebSocket (Edge, default):** `wss://streaming.assemblyai.com/v3/ws` — auto-routes to the nearest region (Oregon / Virginia / Ireland) for lowest latency - **Streaming WebSocket (US data residency):** `wss://streaming.us.assemblyai.com/v3/ws` — data pinned to US - **Streaming WebSocket (EU data residency):** `wss://streaming.eu.assemblyai.com/v3/ws` — data pinned to EU - **LLM Gateway (US):** `https://llm-gateway.assemblyai.com/v1/chat/completions` - **LLM Gateway (EU):** `https://llm-gateway.eu.assemblyai.com/v1/chat/completions` — Claude and Gemini only; OpenAI and Qwen are US-only - **Auth header:** `Authorization: YOUR_API_KEY` (no `Bearer`). Same header is used for REST, streaming WS upgrade, temp-token minting, and LLM Gateway - **Content type:** `application/json` for submit/poll and LLM Gateway; `application/octet-stream` (raw binary) for `/v2/upload` Core REST endpoints: - `POST /v2/upload` — upload a local file (raw binary body, **not multipart**). Returns `{ "upload_url": "..." }`. Max 2.2 GB. - `POST /v2/transcript` — submit a job. Returns transcript object with `id` and `status: "queued"`. Max 5 GB / 10 hours. - `GET /v2/transcript/{id}` — poll. Statuses: `queued`, `processing`, `completed`, `error`. Streaming: - `wss://streaming.assemblyai.com/v3/ws?sample_rate=16000&speech_model=universal-3-5-pro` - `GET https://streaming.assemblyai.com/v3/token?expires_in_seconds=60` — mint a single-use temp token for browser/mobile clients. Optional `max_session_duration_seconds` (60–10800, defaults to 3 h) caps the downstream session length. --- ## 5. `speech_models` semantics `speech_models` on pre-recorded requests is an **ordered fallback list**, not parallel execution. The first model in the array is tried; if it's unavailable (e.g., not yet rolled out to the account, or temporarily unhealthy), the next is used. A single transcript is produced by exactly one model. Recommended default: `["universal-3-5-pro", "universal-2"]` — tries the latest model first, falls back to the stable predecessor. On streaming, the parameter is **singular** (`speech_model=universal-3-5-pro`) — there is no fallback list. Easy to mix up. --- ## 6. Pre-recorded quick start ### SDK (recommended) **Python:** ```python # pip install assemblyai import assemblyai as aai import os aai.settings.api_key = os.environ["ASSEMBLYAI_API_KEY"] config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], # fallback handled by SDK speaker_labels=True, ) transcript = aai.Transcriber(config=config).transcribe("https://assembly.ai/wildfires.mp3") # Or a local path: .transcribe("./recording.wav") if transcript.status == aai.TranscriptStatus.error: raise RuntimeError(transcript.error) print(transcript.text) ``` **Node/JS:** ```javascript // npm install assemblyai import { AssemblyAI } from 'assemblyai'; const client = new AssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY }); const transcript = await client.transcripts.transcribe({ audio: 'https://assembly.ai/wildfires.mp3', // or a local file path / Buffer / stream speech_models: ["universal-3-5-pro", "universal-2"], speaker_labels: true, }); if (transcript.status === 'error') throw new Error(transcript.error); console.log(transcript.text); ``` The SDK handles upload, submit, and polling. You don't need to write the polling loop yourself. ### Raw HTTP (fallback — use only if SDK isn't an option) **Upload a local file** (raw bytes, not multipart): ```bash curl -X POST https://api.assemblyai.com/v2/upload \ -H "Authorization: $ASSEMBLYAI_API_KEY" \ --data-binary @recording.wav # -> { "upload_url": "https://cdn.assemblyai.com/upload/..." } ``` **Submit and poll (Python):** ```python import os, time, requests headers = {"authorization": os.environ["ASSEMBLYAI_API_KEY"]} submit = requests.post( "https://api.assemblyai.com/v2/transcript", headers=headers, json={ "audio_url": "https://assembly.ai/wildfires.mp3", "speech_models": ["universal-3-5-pro", "universal-2"], "speaker_labels": True, }, ) transcript_id = submit.json()["id"] while True: res = requests.get( f"https://api.assemblyai.com/v2/transcript/{transcript_id}", headers=headers, ).json() if res["status"] == "completed": print(res["text"]); break if res["status"] == "error": raise RuntimeError(res["error"]) time.sleep(3) ``` Common optional params: `speaker_labels`, `language_detection`, `language_code`, `punctuate`, `format_text`, `redact_pii`, `redact_pii_audio`, `keyterms_prompt`, `webhook_url`, `prompt`. --- ## 7. Webhooks (skip polling) Provide `webhook_url` on submit; AssemblyAI POSTs when the job finishes: ```json { "transcript_id": "5552493-16d8-42d8-8feb-c2a16b56f6e8", "status": "completed" } ``` Handler requirements: - Return 2xx within **10 seconds**. Otherwise retried up to 10 times, 10s apart. 4xx is not retried. - On receipt, call `GET /v2/transcript/{id}` to fetch the full result — the webhook payload doesn't include it. Optional custom auth on your webhook: set `webhook_auth_header_name` and `webhook_auth_header_value` when submitting. **Source IPs** (for allowlists): US `44.238.19.20`, EU `54.220.25.36`. **Local dev note:** Webhook URLs must be publicly reachable. Use ngrok, Cloudflare Tunnel, or similar during development. --- ## 8. LLM Gateway (chapters, summaries, custom analysis) LLM Gateway replaces both the deprecated transcript params (`auto_chapters`, `summarization`, `summary_model`, `summary_type`) and the legacy **LeMUR** API, which sunset on 2026-03-31. If a developer mentions LeMUR or `transcript_ids`, point them at LLM Gateway and the [migration guide](/llm-gateway/quickstart). Workflow: 1. Transcribe normally with `POST /v2/transcript`. 2. Once `status == "completed"`, POST to LLM Gateway with the transcript text (or paragraphs from `GET /v2/transcript/{id}/paragraphs` for chapter-style output): ```http POST https://llm-gateway.assemblyai.com/v1/chat/completions Authorization: YOUR_API_KEY Content-Type: application/json { "model": "claude-sonnet-4-6", "messages": [ { "role": "system", "content": "Produce a 5-bullet summary of the transcript." }, { "role": "user", "content": "" } ], "max_tokens": 1000 } ``` Model IDs are exact strings — see the [LLM Gateway Overview](/llm-gateway/quickstart) for the current list. Examples: `claude-opus-4-7`, `claude-sonnet-4-6`, `claude-haiku-4-5-20251001`, `gpt-5.2`, `gpt-5.1`, `gpt-4.1`, `gemini-3.5-flash`, `gemini-2.5-pro`, `gemini-2.5-flash`, `qwen3-next-80b-a3b`. `claude-sonnet-4` by itself is **not** valid — always include the version suffix. EU region (`llm-gateway.eu.assemblyai.com`) supports Anthropic and Google only. Do not submit with `auto_chapters` and `summarization` both enabled — the API rejects it (`Only one of the following models can be enabled at a time: auto_chapters, summarization.`). But the broader rule is simpler: **don't use either.** --- ## 9. Streaming — Universal-3.5 Pro **WebSocket (default, Edge Routing):** `wss://streaming.assemblyai.com/v3/ws?sample_rate=16000&speech_model=universal-3-5-pro` For data residency, swap the host: `streaming.us.assemblyai.com` (US-pinned) or `streaming.eu.assemblyai.com` (EU-pinned). The default host auto-routes to the nearest region. **Audio format:** PCM16 signed little-endian, mono, 16 kHz. Binary WebSocket frames, **50–1000 ms per chunk**, no faster than real-time. Phone audio (`encoding=pcm_mulaw`, `sample_rate=8000`) is sent as-is — don't upsample. **Auth:** - Server-side: `Authorization` header on the WS upgrade. - Browser/mobile: mint a short-lived token server-side and pass it as `?token=` (no Authorization header). Mint a token: ```bash curl -s "https://streaming.assemblyai.com/v3/token?expires_in_seconds=60" \ -H "Authorization: $ASSEMBLYAI_API_KEY" # { "token": "..." } ``` `expires_in_seconds` must be 1–600. Tokens are single-use per session. ### Server messages (JSON) - `Begin` — `{ type, id, expires_at }` - `SpeechStarted` — `{ type, timestamp, confidence }` - `Turn` — `{ type, turn_order, end_of_turn, transcript, end_of_turn_confidence, words:[...], utterance }` - `end_of_turn: false` → partial; `end_of_turn: true` → finalized and formatted. Always read `transcript` for current text. - `Termination` — `{ type, audio_duration_seconds, session_duration_seconds }` ### Client messages - Binary PCM16 frames — audio. - `{ "type": "Terminate" }` — graceful end. **Always send this when done.** - `{ "type": "ForceEndpoint" }` — force current turn to end. - `{ "type": "KeepAlive" }` — only needed if `inactivity_timeout` is set. - `{ "type": "UpdateConfiguration", "keyterms_prompt": [...], "min_turn_silence": 100, "max_turn_silence": 1000 }` — adjust mid-session. ### SDK (recommended) **Python:** ```python # pip install "assemblyai>=1.0.0" import os from assemblyai.streaming.v3 import ( StreamingClient, StreamingClientOptions, StreamingEvents, StreamingParameters, TurnEvent, ) def on_turn(_, event: TurnEvent): tag = "FINAL" if event.end_of_turn else "partial" print(f"{tag}: {event.transcript}") client = StreamingClient( StreamingClientOptions(api_key=os.environ["ASSEMBLYAI_API_KEY"]) ) client.on(StreamingEvents.Turn, on_turn) client.connect(StreamingParameters(sample_rate=16000, speech_model="universal-3-5-pro")) # Feed 16 kHz mono PCM16 chunks (50–1000ms each) via client.stream(chunk) # When finished: client.disconnect(terminate=True) # sends Terminate and closes cleanly ``` **Node/JS:** ```javascript // npm install assemblyai import { AssemblyAI } from 'assemblyai'; const client = new AssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY }); const rt = client.streaming.transcriber({ sampleRate: 16000, speechModel: 'universal-3-5-pro', }); rt.on('turn', (turn) => { const tag = turn.end_of_turn ? 'FINAL' : 'partial'; console.log(`${tag}: ${turn.transcript}`); }); rt.on('error', (err) => console.error(err)); await rt.connect(); // rt.sendAudio(pcm16Buffer) for each 50–1000ms chunk // When done: await rt.close(); // sends Terminate and closes ``` ### Raw WebSocket (fallback) **Node.js (`ws`):** ```javascript import WebSocket from 'ws'; const ws = new WebSocket( 'wss://streaming.assemblyai.com/v3/ws?sample_rate=16000&speech_model=universal-3-5-pro', { headers: { authorization: process.env.ASSEMBLYAI_API_KEY } }, ); ws.on('open', () => { // Feed PCM16 16kHz mono chunks here, 50–1000ms each. // Example: audioStream.on('data', (chunk) => ws.send(chunk)); }); ws.on('message', (raw) => { const msg = JSON.parse(raw.toString()); if (msg.type === 'Turn') { console.log(msg.end_of_turn ? `FINAL: ${msg.transcript}` : `partial: ${msg.transcript}`); } }); function stop() { ws.send(JSON.stringify({ type: 'Terminate' })); // required! } ``` **Python (`websockets`):** ```python import asyncio, json, os, websockets URL = "wss://streaming.assemblyai.com/v3/ws?sample_rate=16000&speech_model=universal-3-5-pro" async def run(audio_source): """audio_source: async iterator yielding 50–1000ms PCM16 chunks at 16kHz mono.""" async with websockets.connect( URL, additional_headers={"Authorization": os.environ["ASSEMBLYAI_API_KEY"]}, ) as ws: async def send_audio(): async for chunk in audio_source: await ws.send(chunk) await ws.send(json.dumps({"type": "Terminate"})) async def recv_loop(): async for raw in ws: msg = json.loads(raw) if msg["type"] == "Turn": tag = "FINAL" if msg["end_of_turn"] else "partial" print(f"{tag}: {msg['transcript']}") elif msg["type"] == "Termination": return await asyncio.gather(send_audio(), recv_loop()) # asyncio.run(run(my_audio_iterator())) ``` --- ## 10. Voice Agent API (managed speech-in / speech-out) Use this when the developer wants a complete spoken AI agent — not just transcription. Single WebSocket, audio in and audio out, with STT + LLM + TTS + turn detection + tool calling all managed by AssemblyAI. **Endpoint:** `wss://agents.assemblyai.com/v1/ws` **Auth:** `Authorization: Bearer YOUR_API_KEY` — the Bearer prefix is **required** on this product (different from STT and LLM Gateway, which take the raw key). For browsers/mobile, mint a temp token instead and pass it as `?token=`. **Token endpoint (for browser/mobile clients):** ```bash curl -s "https://agents.assemblyai.com/v1/token?expires_in_seconds=300&max_session_duration_seconds=8640" \ -H "Authorization: Bearer $ASSEMBLYAI_API_KEY" # { "token": "..." } ``` - `expires_in_seconds`: 1–600 (controls how long the token can be redeemed for) - `max_session_duration_seconds`: 60–10800 (caps the resulting session; defaults to the 3-hour max) - Tokens are **single-use** per session — get a fresh one for every reconnect (including `session.resume`). **Audio format:** PCM16 mono **24 kHz**, **base64-encoded inside JSON events** (not raw binary frames — this is different from real-time STT). ~50 ms chunks (2,400 bytes) is fine; the server buffers continuously, exact chunk size doesn't matter. ### Lifecycle (the events that matter) 1. Client connects, sends `session.update` immediately (don't wait for `session.ready`): ```json { "type": "session.update", "session": { "system_prompt": "You are a helpful assistant.", "greeting": "Hi there! How can I help?", "input": { "format": { "encoding": "audio/pcm" }, "keyterms": ["AssemblyAI", "Universal-3"], "turn_detection": { "vad_threshold": 0.5, "min_silence": 200, "max_silence": 1000, "interrupt_response": true } }, "output": { "voice": "anna", "format": { "encoding": "audio/pcm" } }, "tools": [ /* flat-schema tool defs, see step 5 */ ] } } ``` Output `encoding` accepts `audio/pcm` (24 kHz, default), `audio/pcmu` (G.711 μ-law, 8 kHz), or `audio/pcma` (G.711 A-law, 8 kHz) — use the G.711 variants for telephony bridges (Twilio, etc.) so you don't have to resample. 2. Server replies with `session.ready` (capture `session_id` for `session.resume` if you reconnect within 30 s of a disconnect). 3. **Only after `session.ready`**, start streaming mic audio: ```json { "type": "input.audio", "audio": "" } ``` 4. Server emits, in roughly this order, per turn: - `input.speech.started` / `input.speech.stopped` (VAD) - `transcript.user.delta` (partials) and `transcript.user` (final) - `reply.started`, `reply.audio` (multiple base64 PCM16 chunks — write directly into an output buffer at 24 kHz), `transcript.agent`, `reply.done` - **Field-name asymmetry:** `input.audio` carries audio in the `audio` field; `reply.audio` carries it in the `data` field. Easy to miss — copying `event["audio"]` from input handling will silently return nothing on output. 5. **Tool calls:** tool definitions in `session.tools` use a **flat** schema — *not* OpenAI's nested `{type: "function", function: {...}}` form: ```json { "type": "function", "name": "get_weather", "description": "Get the current weather for a city.", "parameters": { "type": "object", "properties": { "location": { "type": "string" } }, "required": ["location"] } } ``` Server sends `tool.call` with `{call_id, name, arguments}`. Accumulate the result locally, then send `tool.result` with the matching `call_id` *after* `reply.done` fires. If `reply.done.status == "interrupted"` (user barge-in), discard pending tool results. 6. **Resume after disconnect:** within 30 s, reconnect with a *new* token and send `session.resume` carrying the previous `session_id` to keep conversation context. After 30 s, start a new session. ### Voices Voice IDs are **exact strings** — invented or remembered values silently fail. Pick from the catalog below or call `GET https://agents.assemblyai.com/v1/voices` for the live list. **English** — `alba`, `eve`, `george`, `jane`, `jean`, `mary`, `michael` (🇺🇸 American), `anna`, `charles`, `paul`, `vera` (🇬🇧 British). **Language-specific** (each speaks the named language plus English): `giovanni` (🇮🇹 Italian), `lola` (🇪🇸 Spanish), `juergen` (🇩🇪 German), `rafael` (🇵🇹 Portuguese), `estelle` (🇫🇷 French). If the developer needs a voice not in this list, *don't* substitute a similar-sounding name — say so and ask. Pre-Voice-Agent-API names like `claire`, `dawn`, `josh`, `grace`, `pete` are **no longer valid** and will be rejected at `session.update`. The legacy voices `ivy`, `james`, `tyler`, `winter`, `bella`, `david`, `kyle`, `helen`, `martha`, `river`, `emma`, `victor`, `eleanor`, `arjun`, `dmitri`, `pierre`, `giulia`, `luca`, `lucia`, `mateo`, and `diego` have been removed — pick a current voice from the catalog above. ### Playback gotcha Don't sleep-schedule audio chunks. Write each `reply.audio` PCM directly to an OS audio buffer (e.g., `sounddevice.OutputStream.write()`) — the OS drains at exactly 24 kHz and absorbs network jitter. Sleep-based timing drifts and produces pops/gaps. On `reply.done.status == "interrupted"`, flush the output buffer (e.g., `speaker.abort(); speaker.start()`) so the user doesn't hear stale agent speech. ### Quickstart pattern (Python sketch) ```python # pip install websockets sounddevice numpy import asyncio, base64, json, os import sounddevice as sd import websockets URL = "wss://agents.assemblyai.com/v1/ws" SAMPLE_RATE = 24_000 async def main(): headers = {"Authorization": f"Bearer {os.environ['ASSEMBLYAI_API_KEY']}"} async with websockets.connect(URL, additional_headers=headers) as ws: await ws.send(json.dumps({ "type": "session.update", "session": { "system_prompt": "You are a helpful assistant.", "greeting": "Hi! How can I help?", "output": {"voice": "anna"}, }, })) ready = asyncio.Event() loop = asyncio.get_running_loop() mic_q: asyncio.Queue = asyncio.Queue() def on_mic(indata, *_): if ready.is_set(): loop.call_soon_threadsafe(mic_q.put_nowait, bytes(indata)) async def pump_mic(): while True: chunk = await mic_q.get() await ws.send(json.dumps({ "type": "input.audio", "audio": base64.b64encode(chunk).decode(), })) with sd.InputStream(samplerate=SAMPLE_RATE, channels=1, dtype="int16", callback=on_mic), \ sd.OutputStream(samplerate=SAMPLE_RATE, channels=1, dtype="int16") as speaker: asyncio.create_task(pump_mic()) async for raw in ws: ev = json.loads(raw) if ev["type"] == "session.ready": ready.set() elif ev["type"] == "reply.audio": import numpy as np speaker.write(np.frombuffer(base64.b64decode(ev["data"]), dtype=np.int16)) elif ev["type"] == "reply.done" and ev.get("status") == "interrupted": speaker.abort(); speaker.start() asyncio.run(main()) ``` For a complete worked example (MCP-tooled agent that talks back), see the [Voice Agent API quickstart](/voice-agents/voice-agent-api). For browser integration, see the [browser integration guide](/voice-agents/voice-agent-api/browser-integration). ### When to choose Voice Agent API vs Real-time STT + your own LLM/TTS - **Voice Agent API (Section 10):** end-to-end conversational agents, fastest to ship, AssemblyAI manages the pipeline. Use when "speech in, speech out" is the whole product. - **Real-time STT + framework (Section 11):** you need a specific LLM, a specific TTS provider, custom turn-detection logic, complex orchestration (LiveKit/Pipecat/Vapi/Vocode/Retell), or features the managed pipeline doesn't expose yet. If they're not sure, ask: *do you want to choose your own LLM and TTS, or is a managed pipeline fine?* That single answer routes them. --- ## 11. Voice Agent framework configs (real-time STT + your own pipeline) This section is for developers who are NOT using the Voice Agent API (Section 10) — they're wiring AssemblyAI Real-time STT into LiveKit, Pipecat, Vapi, Vocode, Retell, or similar, and bringing their own LLM and TTS. The defaults will not be good enough. Common tuning: - **`keyterms_prompt`** — pass proper nouns, product names, and domain terms. For dynamic values (usernames, order IDs), update mid-session via `UpdateConfiguration`. - **Turn silence bounds** — `min_turn_silence` and `max_turn_silence` (ms). Lower values fire end-of-turn faster but risk cutting speakers off. Higher values reduce false finalizations. Form-filling and dictation use cases often want wider windows. - **Multilingual** — Universal-3 Pro Streaming code-switches natively and ignores `end_of_turn_confidence_threshold`. Pass `language_code` to bias the model toward a single language. - **Barge-in / false SpeechStarted** — ambient noise, TTS bleed-through, and PSTN echo can cause spurious `SpeechStarted` events. If the agent is interrupting itself, look here first. Framework-level knobs (e.g., LiveKit's `min_interruption_duration`) often complement, not replace, server-side tuning. - **Phone audio** — 8 kHz mu-law (`pcm_mulaw` at 8000 Hz) should be sent as-is, not upsampled to 16 kHz. Upsampling degrades accuracy. When the developer names one of these frameworks, ask about their specific turn-taking and interruption requirements before defaulting. --- ## 12. Browser patterns **Never put the API key in client code.** ### Pre-recorded — proxy upload + submit through your server ```javascript // Next.js route handler (server) export async function POST(request) { const incoming = await request.formData(); const file = incoming.get('file'); // Blob const upload = await fetch('https://api.assemblyai.com/v2/upload', { method: 'POST', headers: { authorization: process.env.ASSEMBLYAI_API_KEY }, body: file.stream(), duplex: 'half', }); const { upload_url } = await upload.json(); const submit = await fetch('https://api.assemblyai.com/v2/transcript', { method: 'POST', headers: { authorization: process.env.ASSEMBLYAI_API_KEY, 'content-type': 'application/json', }, body: JSON.stringify({ audio_url: upload_url, speech_models: ['universal-3-5-pro', 'universal-2'], }), }); return Response.json(await submit.json()); } ``` ### Streaming — server mints a temp token, client connects directly ```javascript // Server export async function GET() { const res = await fetch( 'https://streaming.assemblyai.com/v3/token?expires_in_seconds=60', { headers: { authorization: process.env.ASSEMBLYAI_API_KEY } }, ); return Response.json(await res.json()); // { token } } ``` ```javascript // Client const { token } = await fetch('/api/aai-token').then((r) => r.json()); const ws = new WebSocket( `wss://streaming.assemblyai.com/v3/ws?sample_rate=16000&speech_model=universal-3-5-pro&token=${token}`, ); ws.onmessage = (e) => { const msg = JSON.parse(e.data); if (msg.type === 'Turn') console.log(msg.transcript, msg.end_of_turn); }; ``` ### Capturing mic audio in the browser `MediaRecorder` does not emit PCM16. You need an `AudioWorklet` (preferred) or `ScriptProcessorNode` to: 1. Capture raw Float32 samples. 2. Downsample to 16 kHz. 3. Convert Float32 → Int16. 4. Send each ~50 ms chunk as a binary WS frame. Reference: [AssemblyAI realtime-transcription-browser-js-example](https://github.com/AssemblyAI/realtime-transcription-browser-js-example). --- ## 13. Not supported / out of scope If a developer asks for any of these, say so directly and propose the closest supported alternative. Do not improvise. - **Real-time translation.** AssemblyAI transcribes, it doesn't translate. Suggest: transcribe with U3 Pro, then translate via LLM Gateway. - **On-device / offline STT.** Cloud API only. - **Speaker identification (matching voices to known people).** `speaker_labels` does diarization (Speaker A, B, C) but does not recognize specific individuals. - **Standalone TTS.** Not an AssemblyAI product *as a separate API*. TTS is bundled into the Voice Agent API (Section 10) — if they need just-TTS, point them to a dedicated provider. - **Voice activity detection as a standalone product.** VAD is internal to the streaming pipeline and surfaced via `SpeechStarted` / turn events, not exposed separately. --- ## 14. Error handling ### REST (pre-recorded) - **401** — Missing/invalid Authorization, disabled account, or insufficient balance. Double-check there's no `Bearer` prefix. - **Transcript `status: "error"`** — Read the `error` field on `GET /v2/transcript/{id}`. - **Retries** — Exponential backoff on 5xx. For 429, respect the `Retry-After` header. - **Limits** — `/v2/upload` max 2.2 GB; `/v2/transcript` max 5 GB / 10 hr per file. - **Scoping** — An API key can only transcribe files uploaded under the same project. ### Streaming — handshake - **HTTP 410** — The old `v2` streaming endpoint is deprecated. Upgrade to `/v3/ws`. This is an HTTP status on the upgrade request, not a WebSocket close code. ### Streaming — WebSocket close codes | Code | Meaning | |------|---------| | `1008` | Unauthorized: missing/invalid Authorization or token | | `3005` | Session cancelled (server-side error) | | `3006` | Invalid message type / invalid JSON | | `3007` | Audio chunk outside 50–1000 ms, or sent faster than real-time | | `3008` | Session expired (3-hour cap) | | `3009` | Too many concurrent sessions | ### Streaming gotchas - `speech_model` (streaming, singular) vs `speech_models` (pre-recorded, plural). Don't mix up. - On U3 Pro Streaming, `end_of_turn_confidence_threshold` is silently ignored; `language_code` biases the multilingual model toward a single language. - Always send `{ "type": "Terminate" }` when finished. An abandoned session stays billable until the 3-hour cap (`3008`). - Chunk size matters: frames outside 50–1000 ms will close the socket with `3007`. --- ## 15. Quick-reference gotchas - No `Bearer` prefix on the Authorization header — *except* for the Voice Agent API (Section 10), which requires `Authorization: Bearer ...`. - `speech_models` is **optional** on pre-recorded submits — it defaults to `["universal-3-5-pro", "universal-2"]` and is an **ordered fallback list** when provided. - `/v2/upload` takes **raw binary**, not multipart. - Webhook handlers must return 2xx in ≤10 seconds. - Local webhook development needs a public tunnel (ngrok, Cloudflare Tunnel). - Browser code never holds the API key. Proxy uploads, or mint temp tokens for streaming. - Always `Terminate` streaming sessions. - Don't use `auto_chapters`, `summarization`, `summary_model`, `summary_type`. Use LLM Gateway. - Medical mode is `domain: "medical-v1"` (pre-recorded body param / streaming query param). The legacy `medical_mode` flag is **not** the right name. - LLM Gateway model IDs are exact and versioned (e.g., `claude-sonnet-4-6`, `gpt-5.2`, `gemini-2.5-pro`). Shorthand like `claude-sonnet-4` is invalid. - Phone audio stays at native 8 kHz mu-law (`encoding=pcm_mulaw`) — don't upsample. - EU customers use `api.eu.assemblyai.com`, `streaming.eu.assemblyai.com`, and `llm-gateway.eu.assemblyai.com`. The default streaming host (`streaming.assemblyai.com`) is **Edge Routing**, not US-pinned — use `streaming.us.assemblyai.com` if you need data residency guarantees on the US side. - Speech-model values are **raw strings** in the SDKs (`"universal-3-5-pro"`, `"universal-2"`). Enum aliases like `aai.SpeechModel.universal_3_pro` do **not** exist — agents that hallucinate them produce code that imports cleanly and fails at runtime. - LeMUR has fully sunset (2026-03-31). Don't generate code that calls LeMUR endpoints or passes `transcript_ids` to a chat-completions API — use LLM Gateway with the transcript text in `messages` instead. ````` The same prompt is published at [agent-instructions.md](/agent-instructions.md) if you'd rather link to it than paste the whole thing. ## Docs MCP server For on-demand lookups during a session — useful when the agent is iterating on a feature and needs to check several pages mid-task — connect the AssemblyAI docs MCP server: ``` https://assemblyai.com/docs/mcp ``` **Claude Code:** ```bash claude mcp add assemblyai-docs --transport http https://assemblyai.com/docs/mcp ``` **Cursor** (`.cursor/mcp.json`): ```json { "mcpServers": { "assemblyai-docs": { "url": "https://assemblyai.com/docs/mcp" } } } ``` Any MCP client that supports [Streamable HTTP transport](https://modelcontextprotocol.io/specification/2025-03-26/basic/transports#streamable-http) can connect. The server exposes tools to search and fetch AssemblyAI documentation pages on demand. ## Claude Code skill The [AssemblyAI skill](https://github.com/AssemblyAI/assemblyai-skill) gives Claude Code curated instructions and context for the Python and JavaScript SDKs, streaming, voice agents, audio intelligence, and more. ```bash claude install-skill https://github.com/AssemblyAI/assemblyai-skill ``` Works with 60+ AI coding agents via the universal skills format: ```bash npx skills add AssemblyAI/assemblyai-skill ``` Verify it's installed: ```bash claude skill list ``` ## Tips for best results - **Be specific about the SDK** — say "use the AssemblyAI Python SDK" or "use the JavaScript SDK" rather than "use AssemblyAI". - **Reference the model** — specify `universal-3-5-pro` for pre-recorded or streaming to avoid outdated model names. - **Set your API key as an environment variable** — `export ASSEMBLYAI_API_KEY=your_key` so the agent can reference it naturally. --- # Models URL: https://www.assemblyai.com/docs/getting-started/models Source: docs/getting-started/models.mdx Navigation: Overview > Getting started Description: AssemblyAI's speech-to-text models and their capabilities AssemblyAI offers several state-of-the-art speech recognition models, each optimized for different use cases. Choose the model that best fits your needs based on accuracy, latency, cost, and language requirements. ## Pre-recorded models
  • Highest accuracy, fastest model
  • Supports 18 languages
  • Native code switching
  • Contextual prompting capabilities
  • Keyterms prompting up to 1,000 words
  • High accuracy, low latency
  • Support across 99 languages
  • Keyterms prompting up to 200 words
  • Code switching
We recommend [Universal-3.5 Pro](/pre-recorded-audio/universal-3-5-pro) for pre-recorded audio transcription. It delivers the highest accuracy and fastest transcription out of the box, with optional contextual prompting support. Universal-3.5 Pro supports 18 languages, for anything outside that set, the system automatically falls back to Universal-2, giving you coverage across 99 languages total without any extra configuration. ## Streaming models
  • Highest accuracy for voice agents
  • Fastest word emissions
  • Advanced prompting capabilities
  • Keyterms prompting up to 100 words
  • 18 languages with native code switching
  • Good balance of speed and cost-effectiveness
  • Multilingual real-time transcription
  • Keyterms prompting up to 100 words
  • 6 languages: en, es, pt, de, fr, it
  • Good balance of speed and cost-effectiveness
  • English transcription
  • Keyterms prompting up to 100 words
  • Intelligent endpointing
We recommend [Universal-3.5 Pro Streaming](/streaming/getting-started/transcribe-streaming-audio) for streaming transcription. It provides the highest accuracy with sub-300ms latency, native multilingual code switching, and advanced prompting support. ## Add-on models Add-on models enhance transcription accuracy for specialized domains. They work alongside your chosen speech model and are billed separately.
  • Improved accuracy for medical terminology
  • Medications, procedures, conditions, and dosages
  • Works with pre-recorded and streaming models
  • 4 languages: en, es, de, fr
### Medical Mode Medical Mode (`domain: "medical-v1"`) is an add-on that enhances transcription accuracy for medical terminology — including medication names, procedures, conditions, and dosages. It is optimized for medical entity recognition to correct terms that other models frequently get wrong. **Supported models:** - Pre-recorded: Universal-3.5 Pro, Universal-2 - Streaming: Universal-3.5 Pro Streaming, Universal-Streaming English, Universal-Streaming Multilingual **Supported languages:** English, Spanish, German, French Medical Mode is billed as a separate add-on. See the [pricing page](https://www.assemblyai.com/pricing) for details. Learn more: [Medical Mode for pre-recorded audio](/pre-recorded-audio/medical-mode) | [Medical Mode for streaming](/streaming/medical-mode) ## Choosing the right model ### Pre-recorded #### Universal-3.5 Pro Universal-3.5 Pro is our most powerful Voice AI model, designed to capture the "hard stuff" that traditional ASR models struggle with. It delivers state-of-the-art accuracy for entities, rare words, and domain-specific terminology out of the box, with code switching and optional prompting for more control. It's also our fastest model, so you get the best accuracy without sacrificing speed. **Best for:** - Applications requiring highest-accuracy transcription - Medical scribes needing clinical grade transcription accuracy - Sales intelligence / Call centers needing native code-switching - Meeting notetakers / recruiting notetakers needing high-quality diarization
**Regional dialects** Universal-3.5 Pro also supports regional dialects and local speech variants out of the box — no special configuration needed. See the full list of [supported dialects](/pre-recorded-audio/supported-languages#regional-dialects-and-variants). [Try Universal-3.5 Pro here](/pre-recorded-audio/universal-3-5-pro) #### Universal-2 Universal-2 offers accurate, cost-effective transcription across 99 languages with low latency. It supports code switching and optional keyterms prompting for domain-specific vocabulary (up to 200 words). Universal-2 is the go-to choice when you need reliable transcription across diverse languages. **Best for:** - High accuracy at lower cost with broad language support - High-volume, price-sensitive batch transcription - Support for over 99 languages - Recommended fallback when a requested language isn't supported by Universal-3.5 Pro
[Try Universal-2 here](/pre-recorded-audio/select-the-speech-model) ### Streaming #### Universal-3.5 Pro Streaming The most accurate model with the fastest word emissions for voice agents that demand the highest quality. Best-in-class accuracy with advanced prompting capabilities, including both [keyterms prompting](/streaming/prompting-and-keyterms) and [native prompting](/streaming/prompting-and-keyterms). Supports English, Spanish, German, French, Portuguese, Italian, Turkish, Dutch, Swedish, Norwegian, Danish, Finnish, Hindi, Vietnamese, Arabic, Hebrew, Japanese, and Mandarin. **Best for:** - Real-time voice agents - Applications requiring premium accuracy - Customer service voice agents needing elite entity accuracy - IVR replacement / binary response detection in short utterances - Agent assist and sales intelligence needing real-time speaker diarization, mid-session dynamic prompting - Multilingual voice agents with native code-switching across 18 languages - Compliance and verbatim recording — disfluency control via prompting
**Regional dialects** Universal-3.5 Pro Streaming also supports regional dialects and local speech variants out of the box, with no special configuration needed. See the full list of [supported dialects](/streaming/getting-started/transcribe-streaming-audio). [Learn more about Universal-3.5 Pro Streaming](/streaming/getting-started/transcribe-streaming-audio) #### Universal-Streaming Multilingual A multilingual transcription model offering a good balance of speed and cost-effectiveness. Supports English, Spanish, German, French, Portuguese, and Italian. Features intelligent endpointing and [keyterms prompting](/streaming/prompting-and-keyterms) support for up to 100 words. **Best for:** - Cost-effective real-time transcription across languages - Cost-sensitive multilingual streaming across EN/ES/DE/FR/PT/IT
[Learn more about Universal-Streaming Multilingual](/streaming/getting-started/transcribe-streaming-audio) #### Universal-Streaming English An English transcription model offering a good balance of speed and cost-effectiveness. Features ~300ms word-by-word immutable transcripts, intelligent endpointing, and [keyterms prompting](/streaming/prompting-and-keyterms) support for up to 100 words. **Best for:** - Cost-effective real-time transcription for English - English-only real-time apps — fastest and cheapest streaming option for English
[Learn more about Universal-Streaming English](/streaming/getting-started/transcribe-streaming-audio) To learn how to specify a model, see [selecting a model for pre-recorded audio](/pre-recorded-audio/select-the-speech-model) or [selecting a model for streaming audio](/streaming/select-the-speech-model). ## Pricing For detailed pricing information, visit our [pricing page](https://www.assemblyai.com/pricing). ### Pre-recorded | Model | Price per Hour | Volume discounts | | --------------- | -------------- | ---------------- | | Universal-3.5 Pro | $0.21/hr | Available | | Universal-2 | $0.15/hr | Available | ### Streaming Streaming is billed per hour of **session duration** — the total time your WebSocket connection stays open — not per hour of audio sent. See [Streaming Speech-to-Text billing](/billing-and-pricing#streaming-speech-to-text-billing) for details. | Model | Price per Hour (session duration) | Volume discounts | | -------------------------------- | --------------------------------- | ---------------- | | Universal-3.5 Pro Streaming | $0.45/hr | Available | | Universal-Streaming Multilingual | $0.15/hr | Available | | Universal-Streaming English | $0.15/hr | Available | For volume discounts, please reach out to sales@assemblyai.com. ## Next steps - Explore [Speech Understanding](/speech-understanding) features like summarization, sentiment analysis, and more - Learn about prompting: [Universal-3.5 Pro prompting guide](/pre-recorded-audio/universal-3-5-pro/prompting) | [Universal-3.5 Pro Streaming prompting guide](/streaming/prompting-and-keyterms) --- # Evaluations URL: https://www.assemblyai.com/docs/evaluations Source: docs/evaluations.mdx Navigation: Overview > Getting started Description: Figure out which STT models are best for your product with an evaluation. Choosing the right Speech-to-text model for your product requires more than reviewing public benchmarks. Public benchmarks can be misleading due to overfitting — models are often trained on the same datasets used for evaluation, inflating their reported accuracy. Running an evaluation on your own audio data is the most reliable way to determine which model performs best for your specific use case. AssemblyAI provides evaluation tools for both pre-recorded and streaming transcription, measuring metrics that matter in production. ## Pre-recorded audio evaluations Assess which pre-recorded audio STT model is best for your use case. Pre-recorded evaluations measure accuracy using metrics like Word Error Rate (WER) and Full-Word Error Rate (FWER), giving you a clear picture of transcription quality on your actual audio. Learn how to evaluate pre-recorded STT models on your own audio data. ## Streaming evaluations Assess which real-time STT model is best for your voice agent or real-time use case. Streaming evaluations focus on latency metrics like Time to First Token (TTFT) and Time to Complete Turn (TTCT) alongside accuracy, since both speed and correctness matter for real-time applications. Learn how to evaluate real-time STT models for voice agents and real-time applications. ## Benchmarks If you want to review AssemblyAI's current model performance before running your own evaluation, see our benchmarks for the latest accuracy and latency numbers: - [Pre-recorded STT benchmarks](/pre-recorded-audio/benchmarks) - [Real-time STT benchmarks](/streaming/benchmarks) --- # Account Management URL: https://www.assemblyai.com/docs/account-management Source: docs/account-management.mdx Navigation: Overview > Getting started Description: Account Management documentation. The [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home) is where you manage everything related to your account: API keys, rate limits, cost and usage, billing, team members, personal settings, and access to the in-browser tools for each API. The dashboard sidebar is organized into two top-level sections: - **APIs** — playgrounds, getting started resources, logs, and comparison tools for Voice Agents, Realtime STT, Pre-Recorded STT, and the LLM Gateway. - **Workspace** — account-wide configuration, split into **Manage** (API keys, rate limits, cost, usage) and **Settings** (personal info, organization & access, members, billing, alerts, activity logs). ## APIs section Each API in the sidebar has its own set of tools. They aren't covered in detail here — see the corresponding product docs — but the options you'll see are: - **Playground** — an in-browser environment for testing the API without writing code. - **Get Started** — a quickstart and links to additional resources for that API. - **Logs** — historical records of past requests (async transcripts, realtime sessions, etc.). - **Compare** — side-by-side comparison of results across models (Pre-Recorded STT). - **Truth File Corrector** — compare an async STT transcript against a ground-truth file (Pre-Recorded STT). ## Workspace > Manage The **Manage** area is where you configure how your account talks to the API. ### API Keys API keys are unique credentials that authenticate requests to the API. Each API key is associated with a specific project, ensuring secure and controlled access. You can create and delete API keys based on your plan: | Usage limits | Free | PAYG | Contracted | Enterprise | | ------------------ | ---- | ---- | ---------- | ---------- | | Number of API keys | 2 | 4 | 25 | Custom | #### Create a new API key 1. Log in to your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home). 2. In the sidebar, go to **Workspace** > **Manage** > **API Keys**. 3. Click **Create New API Key**. 4. Enter a descriptive name (for example, `Production API` or `Development API`). 5. Click **Create**. #### Delete an API key 1. Log in to your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home). 2. In the sidebar, go to **Workspace** > **Manage** > **API Keys**. 3. Locate the API key you want to delete and click **Delete**. 4. Confirm the deletion in the popup dialog. This action cannot be undone. Make sure no active applications are using the key before deletion. Each project must have at least one active API key at all times. You cannot delete the last remaining key in a project. ### Rate limits The **Rate Limits** page shows the request limits in effect for your account. For details on how rate limits work and how to handle them in your application, see the [Pre-Recorded STT rate limits](/pre-recorded-audio/rate-limits) and [Real-time STT rate limits](/streaming/rate-limits) pages. ### Cost and usage The **Cost** and **Usage** pages provide a breakdown of your spend and usage so you can track and manage costs effectively. You can analyze the data at different levels of granularity: - Account - Product (e.g., Speech-to-Text, Streaming, LLM Gateway) - Models (e.g., Best, Claude Sonnet 4.6, etc.) - Project - API key Usage and spend data updates every 2 minutes. For streaming, a session must be closed before its usage populates. ## Workspace > Settings The **Settings** area covers personal account info, organization configuration, team membership, data controls, billing, alerts, and activity logs. ### Personal Info Manage the details tied to your individual user: - Change the email address associated with your account. - Change your password. - Set up or update multi-factor authentication (MFA). - Log out. You can also reach **Personal Info** from anywhere in the dashboard by clicking the portrait icon in the top-right corner. ### Organization & Access Control how members authenticate and what protections are enforced on the organization: - Authentication methods available to members. - MFA enforcement for the organization. - [Single sign-on (SSO)](/sso) configuration. SSO is a paid add-on, billed monthly for each connection. On pay-as-you-go accounts it's paid out of your prepaid balance, so auto-pay has to stay enabled and cover your connections — see [Auto-pay and SSO connections](/billing-and-pricing#auto-pay-and-sso-connections). - **Danger Zone** — delete your account or organization. See [Deleting your account or organization](#deleting-your-account-or-organization) below. ### Members Invite teammates to your account with role-based access control. Each member gets their own login and has a single role per account. A single person can belong to multiple accounts with different roles in each. For example, an Owner of one organization and a Reader on another. Projects, API keys, billing, and all other account configuration are owned by the account, not by individual members. Adding, removing, or changing a member's role has no effect on any of these. #### Roles There are three roles. Each account has exactly one Owner, and a member has exactly one role per account: | Role | Permissions | | ---- | ----------- | | **Owner** | Full control, including billing, account deletion, and ownership transfer. There is exactly one Owner per account. | | **Admin** | Full administrative access — can do everything an Owner can, except delete the account. | | **Reader** | Read-only access. Can view projects, API keys, usage, and transcriptions, but cannot create, edit, or delete. This is the default role for new invites. | Here's a quick reference for what each role can do: | Action | Owner | Admin | Reader | | ------ | ----- | ----- | ------ | | View members, projects, API keys, usage, transcriptions | ✓ | ✓ | ✓ | | Invite, update, or remove members | ✓ | ✓ | ✗ | | Create, edit, or delete projects, tokens, alerts | ✓ | ✓ | ✗ | | View or edit billing, autopay, invoices | ✓ | ✓ | ✗ | | Manage [Data Controls](/data-controls) (opt out of model training, set TTL, sign a BAA) | ✓ | ✓ | ✗ | | Delete the account | ✓ | ✗ | ✗ | Ownership transfer cannot be performed self-serve. Contact [support@assemblyai.com](mailto:support@assemblyai.com) to request a transfer. #### Member limits The number of members (active + pending invites combined) per account depends on your plan: | Plan | Max members | | ---- | ----------- | | Free | 2 | | PAYG | 20 | | Contracted | 20 | | Enterprise | 20 | Pending invites count toward the limit. If you need a higher member limit, contact [support@assemblyai.com](mailto:support@assemblyai.com). #### Inviting a member Owners and Admins can invite team members from the dashboard: 1. Go to **Workspace** > **Settings** > **Members**. 2. Click **Invite Member**. 3. Enter the invitee's email address and select a role. 4. Click **Send Invite**. The invitee receives an email with an accept link. Invite links are valid for 7 days. If one expires, the inviter must send a new one. If your account is at its member limit, revoke a pending invite or remove an existing member before sending a new invitation. #### Revoking an invite To cancel a pending invite, use the delete action on the pending invite entry in the **Members** list. There is no separate revoke action. Deleting handles both pending invites and active members. #### Removing a member Owners and Admins can remove any non-Owner member from the **Members** list. The Owner cannot be removed without first transferring ownership to another member. To transfer ownership, contact [support@assemblyai.com](mailto:support@assemblyai.com). After a transfer, the previous Owner is automatically demoted to Admin. #### MFA enforcement Owners and Admins can require all members to use multi-factor authentication (MFA) from **Workspace** > **Settings** > **Organization & Access**. When enabled, members who haven't set up MFA will be prompted to do so on their next login. The **Members** list shows the MFA enrollment status for each member. ### Data Controls The **Data Controls** page lets Owners and Admins manage how AssemblyAI retains and uses your data — opt out of the model improvement program, set a time-to-live (TTL) for audio and transcripts, and review and sign a Business Associate Agreement (BAA). These controls are available self-serve, at no additional cost, on paid plans. See [Data Controls](/data-controls) for details. ### Billing The **Billing** page shows: - Your current account balance. - Credit card on file (add or update your payment method). - Auto-pay configuration. - Billing history and invoices. You can also reach **Billing** by clicking the **View billing details** button in the bottom-left corner of the dashboard. For details on how billing, auto-pay, and invoices work, see [Billing and Pricing](/billing-and-pricing). ### Alerts Set up usage and balance alerts so you're notified before your balance runs low or your usage crosses a threshold you define. ### Activity Logs The **Activity Logs** page provides an audit trail of account activity — useful for reviewing administrative changes, sign-ins, and other notable events. ## Projects Projects can be used to isolate data for different environments or applications, e.g., production, staging, or development. Each project has its own API keys, allowing for better organization and data access control. Transcripts and other project-specific data are accessible only within the project they were created in — an API key from one project cannot access historical transcripts or data from another project. This separation maintains data security and prevents unintended cross-project access. This project-level scoping also applies to files uploaded via the `/upload` endpoint. An API key can only transcribe files that were uploaded within the same project. If you need to transcribe the same file across multiple projects, you must upload it separately in each project. You can create, rename, and delete projects based on your plan: | Usage limits | Free | PAYG | Contracted | Enterprise | | ------------------ | ---- | ---- | ---------- | ---------- | | Number of projects | 2 | 2 | 5 | Custom | ## Account switching When you sign in, AssemblyAI logs you straight into your account, or shows you a list to choose from if you belong to multiple accounts. If you belong to multiple accounts, an organization switcher is available in the dashboard so you can move between them without logging out. ### Migrating existing accounts to multi-user If you already have multiple AssemblyAI accounts, the recommended path is to designate one as your primary account and invite your other users into it. **A few important notes:** - Existing accounts are **not** automatically merged. Each account retains its own projects, API keys, transcripts, and billing history. - Users invited to another account will retain their original account. They may see multiple accounts in the account switcher after accepting an invite. - Going forward, manage billing changes in a single account to avoid confusion and unexpected charges. - To move workflows, switch the API keys in your applications to point to the primary account's API keys. **Best practice:** Start from the account you want to treat as primary, then invite teammates into that account using the invite flow described above. ## Changing your account email You can change your account email address yourself from **Workspace** > **Settings** > **Personal Info**. You can also reach this page by clicking the portrait icon in the top-right corner of the dashboard. The Owner's email address is used for billing communications. If you [transfer account ownership](#members) to another member, billing emails go to the new Owner's address. ## Deleting your account or organization If you are the sole member of your account, you can delete it from the dashboard: 1. Log in to your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home). 2. In the sidebar, go to **Workspace** > **Settings** > **Organization & Access**. 3. Scroll to the **Danger Zone**. 4. Click the option to delete your account or organization and follow the prompts. Account deletion is permanent and cannot be undone. Make sure to back up any important information before proceeding with deletion. If your account has other members, remove them or transfer ownership before attempting to delete the account. If you encounter any issues, contact [support@assemblyai.com](mailto:support@assemblyai.com). --- # Single Sign-On (SSO) URL: https://www.assemblyai.com/docs/sso Source: docs/sso.mdx Navigation: Overview > Getting started Description: Set up SAML or OIDC Single Sign-On for your AssemblyAI organization, including step-by-step configuration for Okta, Google Workspace, and Microsoft Entra ID. AssemblyAI supports SAML and OIDC Single Sign-On (SSO) for organizations on Contract and PAYG plans. SSO lets your organization members sign in with your company's identity provider (IdP) — such as Okta, Google Workspace, or Microsoft Entra ID — instead of a password or magic link. SSO is self-served: you configure it yourself from [**Settings → Organization & Access**](https://www.assemblyai.com/dashboard/settings/organization-and-access) in the AssemblyAI dashboard. No contact with support is required to get started. Only organization **owners** and **admins** can access and configure SSO and billing settings. With SSO you can: - Let members sign in through your company's IdP - Auto-create new members on their first SSO login with Just-in-Time (JIT) provisioning - Enforce SSO-only login by disabling all other sign-in methods ## Pricing | Detail | Info | | ----------- | -------------------------------------------------------- | | Price | $199/month per connection | | Limit | 1 connection per organization | | Billing | Line item on invoice — review in [**Settings → Billing**](https://www.assemblyai.com/dashboard/settings/billing) | | Eligibility | Contract and PAYG plans | Billing starts when a connection becomes **Active**, not when you create it — a connection sitting in **Pending** while you finish IdP configuration is not charged. SSO billing does not prorate: once a connection goes active, you're charged the full month's price regardless of how many days remain in the billing period. If you delete a connection mid-month, the connection is terminated immediately but the current period's charge is not refunded. ### Balance and auto-pay requirements on PAYG plans On PAYG plans the subscription is charged against your prepaid balance, so two things have to be true before you can add a connection. **Your balance has to cover the connection.** A connection is charged as soon as it becomes active, so you need at least **$199** in balance — one connection's price, whatever else you already have. If you're short when you add a connection, the dashboard offers to charge your card for the difference (plus $10, so a small amount of usage doesn't put you straight back under) and creates the connection once that payment goes through. Without a payment method on file there's nothing to charge, so add one first. **Your auto-pay settings have to keep it covered.** This is what pays for the connection each month after the first, so [auto-pay](/billing-and-pricing#how-does-auto-pay-work) needs: - A refill threshold of at least **$199** for each active connection. - A refill amount at least **$10** above your refill threshold. If your settings are below that, creating or activating a connection raises them for you. It never lowers a threshold or amount you've set higher, and raising them isn't itself a charge — auto-pay charges your card on its own schedule, when your balance falls past the threshold. Both requirements apply when you *create* a connection as well as when it goes active, so you can't end up with a connection you're unable to activate. While a connection is active: - Auto-pay can't be switched off. - Your refill threshold can't drop below the price of your active connections. - If the invoice covering the connection isn't paid, the connection is removed — see [What happens if I delete my SSO connection](#what-happens-if-i-delete-my-sso-connection-while-members-are-using-it) for what members experience when a connection goes away. Deleting the connection lifts these restrictions, and you're free to lower auto-pay again. Contract plans are invoiced at the end of the month and don't use auto-pay, so none of this applies to them. ## How SSO login works **SP-initiated (standard) flow:** 1. A user clicks **Sign in with SSO** on the AssemblyAI login page. 2. The user is redirected to your company's IdP. 3. The IdP authenticates the user and sends the authentication response back to AssemblyAI — a signed SAML assertion for SAML connections, or an authorization code for OIDC connections. 4. AssemblyAI validates the response and logs the user in. **IdP-initiated flow:** 1. A user clicks the AssemblyAI tile in their IdP dashboard (for example, Okta). 2. The IdP sends a SAML assertion directly to AssemblyAI's callback URL. 3. AssemblyAI processes the assertion and logs the user in. ## Set up SSO You configure SSO from [**Settings → Organization & Access**](https://www.assemblyai.com/dashboard/settings/organization-and-access) in the dashboard. When you create a connection, you choose a **connection type**: SAML or OIDC. Which one to pick depends on your IdP — most enterprise IdPs support both, and either works with AssemblyAI. ### SAML setup SAML setup always follows the same pattern, regardless of IdP: 1. **AssemblyAI → IdP**: Copy the **ACS URL** and **Audience URI** from your AssemblyAI SSO connection into your IdP's application settings (sometimes labeled **SP SSO URL** and **SP Entity ID**). 2. **In your IdP**: Configure attribute mapping so the IdP sends the required attributes (see below). 3. **IdP → AssemblyAI**: Copy your IdP's **Metadata URL** (recommended) — or the IdP Entity ID, IdP SSO URL, and X.509 certificate individually — into your AssemblyAI connection. 4. Once all required fields are present and valid, the connection status flips from **Pending** to **Active**. Configuring via **Metadata URL** is recommended wherever your IdP provides one. AssemblyAI auto-populates the Entity ID, SSO URL, and certificate from it, and can re-fetch updated metadata later — for example, after your IdP rotates its signing certificate. ### OIDC setup OIDC setup follows a similar two-direction pattern: 1. **AssemblyAI → IdP**: Copy the **Redirect URL** from your AssemblyAI SSO connection into your IdP as the **Sign-in Redirect URI**, using the **Authorization Code** grant type. 2. **IdP → AssemblyAI**: Copy the **Client ID**, **Client Secret**, and **Issuer URL** (usually your IdP's hostname) from the IdP application into your AssemblyAI connection. 3. Once all required fields are present and valid, the connection status flips from **Pending** to **Active**. The **Issuer URL** must be the plain issuer base URL — not the discovery URL ending in `/.well-known/openid-configuration`. Using the discovery URL is a common cause of failed OIDC setups. ### Required SAML attributes For SAML connections, your IdP must send the following attributes in the SAML assertion: | SAML attribute | Maps to | | -------------- | ---------- | | `email` | Email | | `firstName` | First name | | `lastName` | Last name | Your IdP can send a combined `full_name` attribute instead of separate `firstName` and `lastName` attributes. If these attributes are missing or incorrectly named, login fails with an attribute mapping error. If your IdP lets you configure a **NameID** format, set it to the user's email address. ### JIT provisioning Just-in-Time provisioning controls whether new members are automatically created on their first SSO login: - **Anyone**: Any user who authenticates via SSO is automatically added as a member — no prior invite required. - **Nobody**: Only pre-invited members can log in via SSO. An **Active** connection alone doesn't mean new users can log in. If you want new users auto-created on first login, you also need to set **JIT Provisioning → Anyone**. This is a separate step from the IdP configuration itself. ### Sign-in methods The **Sign-in Methods** setting controls which authentication methods are available to your organization's members: Magic Link, Password, Google OAuth, and SSO. To enforce SSO-only login, disable all non-SSO methods. At least one method must always be enabled. Deleting an SSO connection while SSO-only login is enforced will lock out all members once their sessions expire. Re-enable an alternative sign-in method before deleting the connection. ### MFA Authenticator app (TOTP) is the supported MFA method. Enabling **Require MFA for all members** forces all members to enroll on their next sign-in. If your IdP already enforces MFA, AssemblyAI's MFA requirement may not trigger for SSO users — MFA has already been handled at the IdP level. ## IdP configuration guides The steps below cover the most common identity providers. For IdPs not listed here, follow the generic patterns in [Set up SSO](#set-up-sso) — any SAML 2.0- or OIDC-compliant IdP works. ### Okta (SAML) 1. In the Okta Admin console, go to **Applications → Create App Integration → SAML 2.0**. 2. On the **Configure SAML** screen, enter: - **Single sign-on URL** = your AssemblyAI **ACS URL** - **Audience URI (SP Entity ID)** = your AssemblyAI **Audience URI** - **Name ID format** = EmailAddress - **Application username** = Email - **Attribute Statements**: - `firstName` → `user.firstName` - `lastName` → `user.lastName` - `id` → `user.id` 3. Copy the **Metadata URL** from Okta's **Sign On** settings tab into the Single Sign-On section of your AssemblyAI settings. 4. Assign the app to users under **Assignments** in Okta. If a user isn't assigned to the application in Okta, they can't sign in via SSO — even if JIT provisioning is set to Anyone. If a user exists in Okta but their SSO login fails, check their app assignment first. For more detail, see Okta's [SAML app integration documentation](https://help.okta.com/en-us/content/topics/apps/apps_app_integration_wizard_saml.htm). ### Okta (OIDC) 1. In the Okta Admin console, go to **Applications → Create App Integration → OIDC - OpenID Connect** and select **Web Application** as the application type. 2. Under **General Settings**: - **Grant type** = Authorization Code - **Sign-in redirect URIs** = your AssemblyAI **Redirect URL** 3. Assign the app to users or groups under **Assignments**. 4. Save the app, then copy the following from the **Client Credentials** section into your AssemblyAI SSO connection: - **Client ID** - **Client Secret** - **Issuer URL** = your Okta org URL (for example, `https://your-company.okta.com`) The same assignment requirement applies as with SAML: users who aren't assigned to the application in Okta can't sign in via SSO, even if JIT provisioning is set to Anyone. For more detail, see Okta's [OIDC app integration documentation](https://help.okta.com/en-us/content/topics/apps/apps_app_integration_wizard_oidc.htm). ### Google Workspace (SAML) 1. In the Google Admin console, go to **Apps → Web and mobile apps → Add app → Add custom SAML app**. 2. Google displays an **Option 2** panel with the **IdP Entity ID**, **SSO URL**, and **Certificate**. Paste these into your AssemblyAI SSO connection using the manual configuration option. 3. Back in Google, under **Service provider details**, enter: - **ACS URL** = your AssemblyAI **ACS URL** - **Entity ID** = your AssemblyAI **Audience URI** - **Name ID format** = EMAIL - **Name ID** = Primary email 4. Under **Attributes**, map: - First name → `firstName` - Last name → `lastName` 5. Set **User access** (Service status) to **ON** for the relevant organizational unit or group. Google Workspace doesn't provide a metadata URL for custom SAML apps — the manual configuration option (Entity ID, SSO URL, certificate) is the expected path. This is normal, not a misconfiguration. For more detail, see Google's [custom SAML app documentation](https://support.google.com/a/answer/6087519). ### Microsoft Entra ID / Azure AD (SAML) 1. In the Entra admin center, go to **Enterprise applications → New application → Create your own application → Non-gallery**. 2. Under **Single Sign-On → SAML → Basic SAML Configuration**, enter: - **Identifier (Entity ID)** = your AssemblyAI **Audience URI** - **Reply URL (ACS URL)** = your AssemblyAI **ACS URL** 3. Under **Attributes & Claims**: - Edit the **Unique User Identifier (Name ID)** claim and set its source to `user.primaryauthoritativeemail`. - Add additional claims: `firstName` → `user.givenname`, `lastName` → `user.surname`, `id` → `user.objectid`. 4. Copy the **App Federation Metadata Url** from the **SAML Certificates** section into your AssemblyAI SSO connection. 5. Add users or groups under **Users and groups** in the Entra application. Entra's default Name ID claim is often not the user's email address. If SSO login fails with an "Email format is invalid" error and you're using Entra ID, check the Name ID claim source first — this is the most common Entra misconfiguration. For more detail, see Microsoft's [SAML-based single sign-on documentation](https://learn.microsoft.com/en-us/entra/identity/enterprise-apps/add-application-portal-setup-sso). ### Microsoft Entra ID / Azure AD (OIDC) 1. In the Entra admin center, go to **App registrations → New registration** and select **Accounts in this organizational directory only**. 2. Under **Authentication → Add a platform → Web**, set the **Redirect URI** to your AssemblyAI **Redirect URL**. 3. Under **Certificates & secrets → New client secret**, create a secret and copy the secret **value** into the Client Secret field of your AssemblyAI connection. 4. From the app's **Overview** page: - **Client ID** = the **Application (client) ID** - **Issuer URL** = `https://login.microsoftonline.com//v2.0`, substituting your **Directory (tenant) ID** Two common copy-paste errors to watch for: copy the client secret's **value**, not its secret ID — and build the Issuer URL in the exact format above with your tenant ID substituted in. The Issuer isn't a value you can copy verbatim from a single field in the Entra UI, and it isn't just your Entra domain. For more detail, see Microsoft's [OpenID Connect documentation](https://learn.microsoft.com/en-us/entra/identity-platform/v2-protocols-oidc). ### Attribute name cheat sheet | IdP | Email | First name | Last name | | ------------------ | -------------------------------------------------------- | ----------------------------- | --------------------------- | | Okta | email (via NameID) | `firstName` | `lastName` | | Google Workspace | email (via NameID) | `firstName` | `lastName` | | Microsoft Entra ID | email (via NameID, from `user.primaryauthoritativeemail`) | `firstName` (`user.givenname`) | `lastName` (`user.surname`) | | Auth0 | email | `firstName` | `lastName` | ## Troubleshooting ### I set up SSO but login still goes to the regular login page The most likely causes: 1. **SSO is not enabled as a sign-in method.** Check [**Settings → Organization & Access**](https://www.assemblyai.com/dashboard/settings/organization-and-access) **→ Sign-in Methods** — SSO must be toggled on. 2. **The connection is in Pending status.** The configuration is incomplete. Finish configuring the connection — for SAML, the Entity ID, SSO URL, and certificate must all be present; for OIDC, the Client ID, Client Secret, and Issuer URL. 3. **You're testing an IdP-initiated flow.** Some flows only work SP-initiated. Try initiating login from the AssemblyAI login page instead. 4. **The ACS URL or Audience URI wasn't saved in your IdP.** Copy AssemblyAI's ACS URL and Audience URI into your IdP configuration. ### "SAML Connection attribute mapping is missing either the 'full_name' or both the 'first_name' and 'last_name' fields" Your IdP isn't sending the required name attributes. Configure your IdP to send `firstName` and `lastName` (or a combined `full_name` attribute). This is configured on the IdP side — see the [attribute name cheat sheet](#attribute-name-cheat-sheet) for the IdP-specific field names. ### "Email format is invalid" error Your IdP is sending the email in an unexpected format, or the email attribute mapping points to the wrong field. Ensure your IdP sends the email as a standard email attribute in the SAML assertion. On Microsoft Entra ID, check that the Name ID claim source is set to `user.primaryauthoritativeemail`. ### A user logged in via SSO but wasn't added as a member JIT provisioning is set to **Nobody**. Either change JIT provisioning to **Anyone**, or pre-invite the user from the Members page before they log in via SSO. ### OIDC login fails after setup The most common OIDC misconfigurations: 1. **Redirect URI mismatch.** The Sign-in Redirect URI in your IdP must exactly match the Redirect URL from your AssemblyAI connection. 2. **Wrong Issuer URL.** Use the plain issuer base URL, not the discovery URL ending in `/.well-known/openid-configuration`. 3. **Missing or wrong client secret.** Make sure a client secret is set, and that you copied the secret's value (not its ID). If the secret has expired, generate a new one and update your connection. ### SSO was working, but broke after our IdP certificate rotated If you configured the connection manually (not via Metadata URL), update the X.509 certificate in your AssemblyAI SSO connection with the new certificate. If you used the Metadata URL option, re-save your connection so AssemblyAI re-fetches the updated metadata. If your IdP has multiple signing certificates, make sure the **active** one is configured in your connection. ### What happens if I delete my SSO connection while members are using it? - The connection is deleted immediately. - Members who logged in through the SSO connection being deleted are signed out immediately. Members who signed in through another method (magic link, password, Google OAuth, or a different SSO connection) keep their active sessions. - Future SSO logins will fail on that connection. - If SSO was the only allowed sign-in method, members will be locked out on their next login. Re-enable another sign-in method before deleting the connection. ## Current limitations - Maximum of 1 SSO connection per organization - No SCIM support — user provisioning and deprovisioning from the IdP isn't automated - Custom attribute mapping is limited to email, first name, and last name --- # Billing and Pricing URL: https://www.assemblyai.com/docs/billing-and-pricing Source: docs/billing-and-pricing.mdx Navigation: Overview > Getting started Description: Billing and Pricing documentation. AssemblyAI uses pay-as-you-go pricing with no contracts, minimums, or monthly subscriptions. You only pay for what you use, and failed transcripts aren't charged. For the latest per-hour rates across all models, see the [pricing page](https://www.assemblyai.com/pricing) or the [Models pricing section](/getting-started/models#pricing). ## How billing works - Add a credit card and deposit funds into your account. Funds are drawn down as you use the API. - Rates are listed per hour for simplicity, but pre-recorded audio is pro-rated to the exact second of audio processed. - Credits are deducted only after a successful transcription completes. If a request errors, you aren't charged for it. - Monitor your balance on the **Billing** page (under **Workspace** > **Settings** > **Billing** in the [dashboard](https://www.assemblyai.com/dashboard/home)) and enable [auto-pay](#how-does-auto-pay-work) to avoid service interruptions. End-of-month invoicing is available for enterprise customers. - Set up billing alerts under **Workspace** > **Settings** > **Alerts** in the dashboard to get notified before your balance runs low. ## Payment method and invoices AssemblyAI accepts all major credit cards. ACH transfers are available in some cases — email [support@assemblyai.com](mailto:support@assemblyai.com) to learn more. To update the credit card on your account, log in to the [dashboard](https://www.assemblyai.com/dashboard/home), go to **Workspace** > **Settings** > **Billing**, and update your payment details. You can also click **View billing details** in the bottom-left corner of the dashboard to jump straight to this page. To change the company information shown on your invoices, email [support@assemblyai.com](mailto:support@assemblyai.com) with the details you'd like to appear. ## Pre-recorded Speech-to-Text billing Pre-recorded transcription is billed on the duration of the submitted audio or video file in seconds, multiplied by the hourly rate for the selected speech model. Add-on features (for example, Entity Detection or Medical Mode) are billed separately at their own hourly rates, also pro-rated to the exact second. ## Streaming Speech-to-Text billing **Streaming is billed per session, not per second of audio** Streaming Speech-to-Text is billed on the total duration that your WebSocket connection stays open, not on the amount of audio you send. You're charged for the entire session — including any time the connection is idle with no audio flowing — so it's important to close streaming sessions as soon as you're done with them. Billing for streaming sessions starts when the WebSocket connection opens and stops when the session ends. This model gives you full control over cost: you can keep a stream open continuously for instant response, or open streams on demand to minimize spend. Each open session is billed independently on its own WebSocket-open duration, so concurrent sessions accumulate billed time in parallel. For example, two streaming sessions running simultaneously for 5 minutes each — including a single call that is dual-streamed under two separate session IDs — bill as 10 minutes of total session time at the streaming rate listed in the [Models pricing section](/getting-started/models#pricing). Follow these best practices to avoid unexpected charges: - Always send a session termination message when your application is finished with a stream. This is what stops billing for the session. - If a session is not closed, it will automatically close after 3 hours — and you'll be billed for the full 3 hours of session time, regardless of how much audio was actually streamed. - Improperly closed streaming sessions are the most common cause of unexpected charges that lead to negative account balances. For details on how to terminate a session, see the [Session termination](/streaming/getting-started/transcribe-streaming-audio) message sequence reference. ## Voice Agent API billing **Voice Agent sessions are billed on WebSocket-open duration** Voice Agent sessions are billed on the total time the WebSocket connection stays open, not on how much audio the user or agent speaks. Always send [`session.end`](/voice-agents/voice-agent-api/events-reference#sessionend) before closing the socket on an intentional disconnect. Billing for a Voice Agent session starts when the WebSocket connection opens and stops when the session ends. Sessions bill at the Voice Agent rate listed on the [pricing page](https://www.assemblyai.com/pricing), pro-rated to the second, and each concurrent session accumulates billed time independently. A few Voice Agent-specific behaviors to be aware of: * **30-second resume grace window.** If the client closes the WebSocket without sending `session.end`, the server keeps the session open for 30 seconds so the client can reconnect with [`session.resume`](/voice-agents/voice-agent-api/events-reference#sessionresume). That grace window is billable. Send `session.end` on any intentional disconnect (user hung up, "End call" button, page unload) to avoid it. See [Unexpected billing after the call ended](/voice-agents/voice-agent-api/troubleshooting#unexpected-billing-after-the-call-ended). * **Idle time is billable.** Just like Streaming STT, silence and idle periods on an open Voice Agent connection are billed. Close sessions as soon as you're done with them. * **Any LLM Gateway calls made by the agent are billed separately** at the LLM Gateway token rates for the model in use — see [LLM Gateway billing](#llm-gateway-billing) below. For the full session lifecycle and best practices, see the [Voice Agent API overview](/voice-agents/voice-agent-api) and [Ending the session cleanly](/voice-agents/voice-agent-api/browser-integration#6-ending-the-session-cleanly). ## LLM Gateway billing LLM Gateway is billed on **input and output tokens**, not on audio duration or session time. Each request is charged at the per-1M-token input and output rates for the specific model you call. Rates vary by model — Claude, GPT, Gemini, and the other supported providers each have their own input and output prices, listed in the Rates table on the [Billing page](https://www.assemblyai.com/dashboard/home) of the dashboard and on the [pricing page](https://www.assemblyai.com/pricing). A few things to keep in mind: * **Output tokens usually cost more than input tokens.** Output pricing reflects the compute needed to generate the response. See [Understanding input and output tokens for LLM Gateway](/faq/understanding-input-and-output-tokens-for-llm-gateway). * **LLM Gateway is not covered by the \$50 free tier.** The free credits granted to new accounts apply to Pre-recorded STT, Real-time STT, Voice Agent API, Speech Understanding, and Guardrails. LLM Gateway usage is billed from your account balance from the first request. * **In-region endpoints have a 10% surcharge as of July 1, 2026.** Requests to the US (`llm-gateway.assemblyai.com`) or EU (`llm-gateway.eu.assemblyai.com`) in-region endpoints are 10% higher than global routing as a direct pass-through of provider price increases, with no AssemblyAI upcharge. To keep the standard rate, opt into [global routing](/llm-gateway/cloud-endpoints-and-data-residency#global-routing). * **Streaming and Voice Agent calls that invoke LLM Gateway are billed additively.** The Streaming or Voice Agent session is billed on its WebSocket-open duration, and any LLM Gateway request made during that session is billed separately on its input and output tokens. To estimate input token cost ahead of time, see [Estimate input token costs for LLM Gateway](/guides/counting-tokens). ## Multichannel billing When [multichannel transcription](/pre-recorded-audio/transcribe-multiple-audio-channels) is enabled, each channel is transcribed and billed separately. The total cost is the audio duration multiplied by the model's hourly rate, multiplied by the number of channels. For example, a 5-minute recording with three channels is billed as 15 minutes of audio (5 minutes × 3 channels) at the selected model's rate. ## Free tier and credits New accounts receive \$50 in free credits for Pre-recorded STT, Real-time STT, Voice Agent API, Speech Understanding, and Guardrails. Credits do not expire, and any unused credits are retained on your account when you upgrade by adding a credit card. LLM Gateway is not included in the free tier. Once your free credits are used up, add a credit card to keep using the API. If your balance reaches \$0 without [auto-pay](#how-does-auto-pay-work) enabled, API access is paused until you top up. ## Volume discounts If you plan to send large volumes of audio or video through the API, [contact sales](https://www.assemblyai.com/contact) to see if you qualify for a volume discount. ## Startup and Y Combinator pricing Early-stage startups can apply for the [AssemblyAI Startup Program](https://www.assemblyai.com/startup-program), designed to help startups build with the speech-to-text API without financial constraints. Y Combinator companies qualify for special pricing — [contact sales](https://www.assemblyai.com/contact) to discuss the discount and see if you qualify. ## AWS Marketplace AssemblyAI is available on the [AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-n24l7xlhzr4o6). Purchasing there lets you consolidate billing with your existing AWS account and apply your spend toward AWS committed-use agreements. Marketplace changes billing only. Your usage still runs on the AssemblyAI cloud API. Endpoints, API keys, and code are unchanged, and there is nothing to migrate in your integration. ### Subscribing 1. Subscribe through the [listing](https://aws.amazon.com/marketplace/pp/prodview-n24l7xlhzr4o6). New subscriptions include a 90-day trial with \$50 in usage credits, then enroll automatically in pay-as-you-go. 2. Complete the fulfillment step, which links your AWS entitlement to a single AssemblyAI workspace. **The entitlement attaches to one workspace** If your organization has more than one AssemblyAI account, the entitlement links to whichever workspace completes the fulfillment step. Accounts are never merged, so billing will not follow usage on a different account. Before subscribing, decide which workspace should be the billed one and make sure the person subscribing on the AWS side signs in to that workspace during fulfillment. See [Migrating existing accounts to multi-user](/account-management#migrating-existing-accounts-to-multi-user). ### If your plan still shows Free after subscribing 1. Check the subscription start date in your AWS console. A future-dated subscription means the entitlement is not effective yet, and re-running the fulfillment link will not change your plan. 2. Confirm which workspace the entitlement resolved to. Your AWS account team can confirm whether the entitlement was issued and where it was mapped. 3. To avoid being blocked in the meantime, add a card under **Workspace** > **Settings** > **Billing**. This converts the workspace to pay-as-you-go immediately, your existing credits are retained and drawn down first, and billing consolidates back to AWS once the entitlement lands. ### Discounted rates The listing self-serves at list pricing. For volume-based rates, [contact sales](https://www.assemblyai.com/contact/sales) to request a private offer, which you accept through AWS. Marketplace billing is tied to an offer with a term. When the offer expires, a new one must be accepted to keep billing running through AWS. ## Billing and customer usage tracking The **Cost** and **Usage** pages under **Workspace** > **Manage** in your dashboard provide a breakdown of spend and usage. If you need to track usage programmatically on a per-customer basis, you can use webhooks with custom query parameters for async transcription, or capture the `session_duration_seconds` from the WebSocket Termination event for streaming. For async transcription, append a `customer_id` query parameter to your webhook URL when submitting a transcription: ``` https://your-domain.com/webhook?customer_id=customer_123 ``` When the transcription completes, extract the `customer_id` from the webhook URL and retrieve the `audio_duration` from the transcript response to record usage. For streaming transcription, manage customer IDs in your application state and capture the `session_duration_seconds` from the Termination event. AssemblyAI bills streaming based on session duration. Creating separate API keys for each customer is not recommended. Use webhooks with metadata instead. ## How does Auto-pay work? Auto-pay automatically recharges your account when your balance falls below a specified threshold. When triggered, it charges your card to bring your balance back to a predetermined amount. Auto-pay is recommended for production environments to prevent service interruptions. ### Example If you set: - **Whenever my balance falls below:** \$25 - **Bring my balance back to:** \$50 When your balance of \$26 drops to \$21 after a \$5 charge, auto-pay will add \$29 to reach your \$50 target balance. When your balance hits \$0 or goes negative without auto-pay, your API access is restricted until you add funds. You'll receive this error message: `{"error": "Your current account balance is negative. Please top up to continue using the API."}` Enable auto-pay and maintain a healthy balance to ensure uninterrupted API access. Set your threshold based on your typical monthly usage to avoid frequent small charges. You can also disable auto-pay at any time from **Workspace** > **Settings** > **Billing** in the [dashboard](https://www.assemblyai.com/dashboard/home), unless your account has an active [single sign-on connection](#auto-pay-and-sso-connections). ### Auto-pay and SSO connections [Single sign-on](/sso) is a paid add-on, billed monthly for each connection at \$199 per connection. On pay-as-you-go accounts that subscription is charged against your prepaid balance, which means two things. **Your balance has to cover a connection before you can add one** — at least \$199, one connection's price. If you're short, the dashboard offers to charge your card for the difference and adds the connection once that payment clears. Without a payment method on file, add one first. **Auto-pay is what keeps it covered after that**, so your settings have to be at least: - **Whenever my balance falls below:** \$199 for each active connection. - **Bring my balance back to:** \$10 above your threshold. Creating or activating a connection raises your auto-pay settings to that minimum if they're below it. It never lowers settings you've already chosen, and raising them isn't itself a charge — auto-pay charges your card on its own schedule. While a connection is active, auto-pay can't be switched off and your threshold can't drop below the price of your active connections. If the invoice covering a connection isn't paid, the connection is removed. Deleting the connection releases the requirement, so you're free to lower auto-pay again afterwards. End-of-month invoiced accounts are billed in arrears and don't use auto-pay, so these requirements don't apply. --- # Introducing Universal-3.5 Pro URL: https://www.assemblyai.com/docs/getting-started/universal-3-5-pro Source: docs/getting-started/universal-3-5-pro.mdx Navigation: Overview > Getting started Description: Learn how to transcribe audio using Universal-3.5 Pro. ## Overview Universal-3.5 Pro is our most powerful Voice AI model, designed to capture the "hard stuff" that traditional ASR models struggle with. It delivers state-of-the-art accuracy for entities, rare words, and domain-specific terminology out of the box, with code switching and optional prompting for more control. It's also our fastest model, so you get the best accuracy without sacrificing speed. Universal-3.5 Pro is available for both **pre-recorded (async)** and **streaming** use cases. Configuration and settings differ between the two because streaming is optimized for real-time audio utterances typically under 10 seconds, with special efficiencies built into the model for low-latency turn detection and voice agent workflows. Based on your use case, navigate to the appropriate guide below: For pre-recorded audio files. Supports long-form audio, prompting, keyterms prompting, and full language detection. For real-time audio streams. Optimized for low-latency turn detection, voice agents, and live transcription. --- # End-to-end examples URL: https://www.assemblyai.com/docs/getting-started/end-to-end-examples Source: docs/getting-started/end-to-end-examples.mdx Navigation: Overview > Getting started > End-to-end examples Description: Copy-paste pipelines that combine multiple AssemblyAI products in a single script. ## Overview Each example below is a self-contained script that wires together several AssemblyAI products into a working pipeline. Run one, see the polished output, and customize from there. Every example uses placeholder API keys (`YOUR_API_KEY`). Replace them with your actual key from the [AssemblyAI dashboard](https://www.assemblyai.com/dashboard). --- ## Pre-recorded pipelines These pipelines transcribe an existing audio file, then enrich the transcript with Speech Understanding features and LLM Gateway analysis. | Pipeline | Products used | Best for | | --- | --- | --- | | [Meeting notetaker](/getting-started/end-to-end-examples/meeting-notetaker) | STT + speaker diarization + Speaker Identification + language detection + LLM Gateway | Team meetings, standups, all-hands | | [Sales call intelligence](/getting-started/end-to-end-examples/sales-call-intelligence) | STT + speaker diarization + Speaker Identification + sentiment analysis + LLM Gateway | Revenue teams, coaching, QA | | [Medical scribe](/getting-started/end-to-end-examples/medical-scribe) | STT + speaker diarization + Speaker Identification + Medical Mode + entity detection + LLM Gateway | Clinical documentation, SOAP notes | | [Content repurposing](/getting-started/end-to-end-examples/content-repurposing) | STT + key phrases + LLM Gateway | Podcasts, webinars, marketing | --- ## Streaming pipelines These pipelines use the [Real-time STT API](/streaming/getting-started/transcribe-streaming-audio) to transcribe audio in real time from a microphone, with optional LLM Gateway integration for live analysis. | Pipeline | Products used | Best for | | --- | --- | --- | | [Real-time meeting assistant](/getting-started/end-to-end-examples/real-time-meeting-assistant) | Real-time STT + Universal-3.5 Pro + LLM Gateway | Live captions, real-time summaries | | [Real-time live captioner](/getting-started/end-to-end-examples/real-time-live-captioner) | Real-time STT + keyterms prompting | Accessibility, live events | --- ## Customize and extend Each pipeline above is a starting point. Here are common ways to build on them: - **Swap LLM models** — Change the `model` parameter in LLM Gateway requests to use any of the [25+ supported models](/llm-gateway/available-models) (Claude, GPT, Gemini, and more). - **Add structured output** — Use [Structured Outputs](/llm-gateway/structured-outputs) to constrain LLM responses to a JSON schema for easier downstream processing. - **Add PII redaction** — Enable [PII Redaction](/guardrails/redact-pii-from-transcripts) to automatically mask sensitive information before it reaches the LLM. - **Use Speaker Identification** — Replace generic speaker labels with real names using [Speaker Identification](/speech-understanding/speaker-identification). - **Add Translation** — Translate transcripts into 86 languages using [Translation](/speech-understanding/translation). - **Use webhooks** — Replace polling with [webhooks](/pre-recorded-audio/webhooks) for production workloads so your server gets notified when transcription completes. ## Next steps - [Pre-recorded STT quickstart](/pre-recorded-audio/getting-started/transcribe-an-audio-file) — Step-by-step guide for your first transcription - [Real-time STT quickstart](/streaming/getting-started/transcribe-streaming-audio) — Set up real-time transcription - [LLM Gateway overview](/llm-gateway/quickstart) — Explore all available models and features --- # Meeting notetaker URL: https://www.assemblyai.com/docs/getting-started/end-to-end-examples/meeting-notetaker Source: docs/getting-started/end-to-end-examples/meeting-notetaker.mdx Navigation: Overview > Getting started > End-to-end examples Description: Transcribe a team meeting with speaker labels, identify speakers by name, then generate structured meeting notes with LLM Gateway. Transcribe a meeting recording with speaker labels and automatic language detection, identify speakers by name, then send the transcript to LLM Gateway for a formatted summary with action items. **Products used:** [Pre-recorded STT](/pre-recorded-audio/getting-started/transcribe-an-audio-file) + [speaker diarization](/pre-recorded-audio/label-speakers) + [Speaker Identification](/speech-understanding/speaker-identification) + [language detection](/pre-recorded-audio/language-detection) + [LLM Gateway](/llm-gateway/quickstart) **Model selection:** This example uses both `universal-3-5-pro` and `universal-2` for broad language coverage across 99 languages. If your meetings are English-only, you can use `universal-3-5-pro` alone for the highest accuracy. ```python expandable import requests import time # ── Config ──────────────────────────────────────────────────── base_url = "https://api.assemblyai.com" headers = {"authorization": "YOUR_API_KEY"} audio_url = "https://assembly.ai/wildfires.mp3" # ── Step 1: Transcribe with speaker labels + language detection ── data = { "audio_url": audio_url, "language_detection": True, "speaker_labels": True, } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) response.raise_for_status() transcript_id = response.json()["id"] while True: result = requests.get(f"{base_url}/v2/transcript/{transcript_id}", headers=headers).json() if result["status"] == "completed": break elif result["status"] == "error": raise RuntimeError(f"Transcription failed: {result['error']}") time.sleep(3) # ── Step 2: Identify speakers by name ── understanding_response = requests.post( "https://llm-gateway.assemblyai.com/v1/understanding", headers=headers, json={ "transcript_id": transcript_id, "speech_understanding": { "request": { "speaker_identification": { "speaker_type": "name", "known_values": ["Alice", "Bob"], # Replace with actual participant names } } }, }, ) understanding_response.raise_for_status() identified = understanding_response.json() # ── Step 3: Format identified transcript for the LLM ── speaker_transcript = "\n".join( f"{u['speaker']}: {u['text']}" for u in identified["utterances"] ) # ── Step 4: Generate meeting notes via LLM Gateway ── llm_response = requests.post( "https://llm-gateway.assemblyai.com/v1/chat/completions", headers=headers, json={ "model": "claude-sonnet-4-6", "messages": [ { "role": "user", "content": ( "You are a meeting notes assistant. Given the transcript below, produce:\n" "1. A concise summary (3-5 sentences)\n" "2. Key decisions made\n" "3. Action items with owners (use speaker labels)\n\n" f"Transcript:\n{speaker_transcript}" ), } ], "max_tokens": 2000, }, ) llm_response.raise_for_status() print("=== Meeting Notes ===\n") print(llm_response.json()["choices"][0]["message"]["content"]) ``` ```javascript expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "YOUR_API_KEY", "Content-Type": "application/json", }; const audioUrl = "https://assembly.ai/wildfires.mp3"; // Step 1: Transcribe with speaker labels + language detection let res = await fetch(`${baseUrl}/v2/transcript`, { method: "POST", headers, body: JSON.stringify({ audio_url: audioUrl, language_detection: true, speaker_labels: true, }), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const { id: transcriptId } = await res.json(); let result; while (true) { res = await fetch(`${baseUrl}/v2/transcript/${transcriptId}`, { headers }); result = await res.json(); if (result.status === "completed") break; if (result.status === "error") throw new Error(`Transcription failed: ${result.error}`); await new Promise((r) => setTimeout(r, 3000)); } // Step 2: Identify speakers by name const understandingRes = await fetch( "https://llm-gateway.assemblyai.com/v1/understanding", { method: "POST", headers, body: JSON.stringify({ transcript_id: transcriptId, speech_understanding: { request: { speaker_identification: { speaker_type: "name", known_values: ["Alice", "Bob"], // Replace with actual participant names }, }, }, }), } ); if (!understandingRes.ok) throw new Error(`Error: ${understandingRes.status}`); const identified = await understandingRes.json(); // Step 3: Format identified transcript for the LLM const speakerTranscript = identified.utterances .map((u) => `${u.speaker}: ${u.text}`) .join("\n"); // Step 4: Generate meeting notes via LLM Gateway res = await fetch("https://llm-gateway.assemblyai.com/v1/chat/completions", { method: "POST", headers, body: JSON.stringify({ model: "claude-sonnet-4-6", messages: [ { role: "user", content: "You are a meeting notes assistant. Given the transcript below, produce:\n" + "1. A concise summary (3-5 sentences)\n" + "2. Key decisions made\n" + "3. Action items with owners (use speaker labels)\n\n" + `Transcript:\n${speakerTranscript}`, }, ], max_tokens: 2000, }), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const llmResult = await res.json(); console.log("=== Meeting Notes ===\n"); console.log(llmResult.choices[0].message.content); ``` ```text === Meeting Notes === ## Summary The discussion covered the impact of Canadian wildfire smoke on US air quality. Experts explained how particulate matter affects respiratory and cardiovascular health. The group reviewed current air quality index readings and discussed protective measures for affected communities. ## Key decisions - Monitor AQI levels daily until smoke clears - Issue public health advisories for sensitive groups ## Action items - Alice: Compile daily AQI data for the affected regions - Bob: Draft public advisory messaging for distribution - Alice: Coordinate with local health departments on response protocols ``` Speaker Identification maps generic labels like "Speaker A" to real names. You can pass a list of `known_values` to guide identification, or omit it to let the model infer names from the conversation. Learn more in the [Speaker Identification guide](/speech-understanding/speaker-identification). --- See the [End-to-end examples overview](/getting-started/end-to-end-examples) for all available pipelines. --- # Sales call intelligence URL: https://www.assemblyai.com/docs/getting-started/end-to-end-examples/sales-call-intelligence Source: docs/getting-started/end-to-end-examples/sales-call-intelligence.mdx Navigation: Overview > Getting started > End-to-end examples Description: Transcribe a sales call with speaker labels and sentiment analysis, identify speakers by role, then generate a coaching scorecard with LLM Gateway. Transcribe a sales call with speaker labels and sentiment analysis, identify speakers by role, then use LLM Gateway to generate a coaching scorecard with talk/listen ratio and sentiment insights. **Products used:** [Pre-recorded STT](/pre-recorded-audio/getting-started/transcribe-an-audio-file) + [speaker diarization](/pre-recorded-audio/label-speakers) + [Speaker Identification](/speech-understanding/speaker-identification) + [sentiment analysis](/speech-understanding/sentiment-analysis) + [LLM Gateway](/llm-gateway/quickstart) **Model selection:** Uses `universal-3-5-pro` for the highest English accuracy. For multilingual sales teams, add `universal-2` as a fallback. ```python expandable import requests import time from collections import Counter # ── Config ──────────────────────────────────────────────────── base_url = "https://api.assemblyai.com" headers = {"authorization": "YOUR_API_KEY"} audio_url = "https://assembly.ai/wildfires.mp3" # ── Step 1: Transcribe with speaker labels + sentiment analysis ── data = { "audio_url": audio_url, "speaker_labels": True, "sentiment_analysis": True, } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) response.raise_for_status() transcript_id = response.json()["id"] while True: result = requests.get(f"{base_url}/v2/transcript/{transcript_id}", headers=headers).json() if result["status"] == "completed": break elif result["status"] == "error": raise RuntimeError(f"Transcription failed: {result['error']}") time.sleep(3) # ── Step 2: Identify speakers by role ── understanding_response = requests.post( "https://llm-gateway.assemblyai.com/v1/understanding", headers=headers, json={ "transcript_id": transcript_id, "speech_understanding": { "request": { "speaker_identification": { "speaker_type": "role", "known_values": ["Sales Rep", "Customer"], } } }, }, ) understanding_response.raise_for_status() identified = understanding_response.json() # ── Step 3: Calculate talk/listen ratio per speaker ── speaker_durations = Counter() for utterance in identified["utterances"]: duration_ms = utterance["end"] - utterance["start"] speaker_durations[utterance["speaker"]] += duration_ms total_ms = sum(speaker_durations.values()) talk_ratios = { speaker: round(dur / total_ms * 100, 1) for speaker, dur in speaker_durations.items() } # ── Step 4: Summarize sentiment shifts ── sentiment_by_speaker = {} for s in result["sentiment_analysis_results"]: speaker = s.get("speaker", "Unknown") sentiment_by_speaker.setdefault(speaker, []).append(s["sentiment"]) sentiment_summary = "" for speaker, sentiments in sentiment_by_speaker.items(): counts = Counter(sentiments) sentiment_summary += ( f"{speaker}: " f"{counts.get('POSITIVE', 0)} positive, " f"{counts.get('NEUTRAL', 0)} neutral, " f"{counts.get('NEGATIVE', 0)} negative\n" ) # ── Step 5: Format transcript and generate coaching scorecard ── speaker_transcript = "\n".join( f"{u['speaker']}: {u['text']}" for u in identified["utterances"] ) llm_response = requests.post( "https://llm-gateway.assemblyai.com/v1/chat/completions", headers=headers, json={ "model": "claude-sonnet-4-6", "messages": [ { "role": "user", "content": ( "You are a sales coaching assistant. Analyze this sales call and produce a scorecard.\n\n" f"Talk/listen ratios: {talk_ratios}\n\n" f"Sentiment breakdown:\n{sentiment_summary}\n" f"Transcript:\n{speaker_transcript}\n\n" "Produce:\n" "1. Call summary (2-3 sentences)\n" "2. Talk/listen ratio analysis (ideal is 40/60 for the rep)\n" "3. Customer sentiment shifts and what caused them\n" "4. Top 3 coaching suggestions for the sales rep" ), } ], "max_tokens": 2000, }, ) llm_response.raise_for_status() print("=== Sales Call Scorecard ===\n") print(llm_response.json()["choices"][0]["message"]["content"]) ``` ```javascript expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "YOUR_API_KEY", "Content-Type": "application/json", }; const audioUrl = "https://assembly.ai/wildfires.mp3"; // Step 1: Transcribe with speaker labels + sentiment analysis let res = await fetch(`${baseUrl}/v2/transcript`, { method: "POST", headers, body: JSON.stringify({ audio_url: audioUrl, speaker_labels: true, sentiment_analysis: true, }), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const { id: transcriptId } = await res.json(); let result; while (true) { res = await fetch(`${baseUrl}/v2/transcript/${transcriptId}`, { headers }); result = await res.json(); if (result.status === "completed") break; if (result.status === "error") throw new Error(`Transcription failed: ${result.error}`); await new Promise((r) => setTimeout(r, 3000)); } // Step 2: Identify speakers by role const understandingRes = await fetch( "https://llm-gateway.assemblyai.com/v1/understanding", { method: "POST", headers, body: JSON.stringify({ transcript_id: transcriptId, speech_understanding: { request: { speaker_identification: { speaker_type: "role", known_values: ["Sales Rep", "Customer"], }, }, }, }), } ); if (!understandingRes.ok) throw new Error(`Error: ${understandingRes.status}`); const identified = await understandingRes.json(); // Step 3: Calculate talk/listen ratio per speaker const speakerDurations = {}; for (const u of identified.utterances) { speakerDurations[u.speaker] = (speakerDurations[u.speaker] || 0) + (u.end - u.start); } const totalMs = Object.values(speakerDurations).reduce((a, b) => a + b, 0); const talkRatios = {}; for (const [speaker, dur] of Object.entries(speakerDurations)) { talkRatios[speaker] = ((dur / totalMs) * 100).toFixed(1) + "%"; } // Step 4: Summarize sentiment shifts const sentimentBySpeaker = {}; for (const s of result.sentiment_analysis_results) { const speaker = s.speaker || "Unknown"; if (!sentimentBySpeaker[speaker]) sentimentBySpeaker[speaker] = []; sentimentBySpeaker[speaker].push(s.sentiment); } let sentimentSummary = ""; for (const [speaker, sentiments] of Object.entries(sentimentBySpeaker)) { const counts = { POSITIVE: 0, NEUTRAL: 0, NEGATIVE: 0 }; sentiments.forEach((s) => counts[s]++); sentimentSummary += `${speaker}: ${counts.POSITIVE} positive, ${counts.NEUTRAL} neutral, ${counts.NEGATIVE} negative\n`; } // Step 5: Format transcript and generate coaching scorecard const speakerTranscript = identified.utterances .map((u) => `${u.speaker}: ${u.text}`) .join("\n"); res = await fetch("https://llm-gateway.assemblyai.com/v1/chat/completions", { method: "POST", headers, body: JSON.stringify({ model: "claude-sonnet-4-6", messages: [ { role: "user", content: "You are a sales coaching assistant. Analyze this sales call and produce a scorecard.\n\n" + `Talk/listen ratios: ${JSON.stringify(talkRatios)}\n\n` + `Sentiment breakdown:\n${sentimentSummary}\n` + `Transcript:\n${speakerTranscript}\n\n` + "Produce:\n" + "1. Call summary (2-3 sentences)\n" + "2. Talk/listen ratio analysis (ideal is 40/60 for the rep)\n" + "3. Customer sentiment shifts and what caused them\n" + "4. Top 3 coaching suggestions for the sales rep", }, ], max_tokens: 2000, }), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const llmResult = await res.json(); console.log("=== Sales Call Scorecard ===\n"); console.log(llmResult.choices[0].message.content); ``` ```text expandable === Sales Call Scorecard === ## Call summary This call discussed the environmental and health impacts of wildfire smoke on US communities. The speakers covered air quality data, health risks, and recommended precautions for the public. ## Talk/listen ratio Sales Rep: 65.3% | Customer: 34.7% Analysis: The ratio is inverted from the ideal 40/60 split. The rep dominated the conversation — focus on asking more open-ended questions. ## Customer sentiment shifts - Started neutral during introductions - Shifted negative when discussing health risks and poor air quality readings - Returned to neutral during the action-planning portion ## Coaching suggestions 1. Ask more discovery questions early to understand the customer's specific concerns 2. When the customer expresses concern, acknowledge before pivoting to solutions 3. Summarize key points at the end and confirm next steps with clear ownership ``` --- See the [End-to-end examples overview](/getting-started/end-to-end-examples) for all available pipelines. --- # Medical scribe URL: https://www.assemblyai.com/docs/getting-started/end-to-end-examples/medical-scribe Source: docs/getting-started/end-to-end-examples/medical-scribe.mdx Navigation: Overview > Getting started > End-to-end examples Description: Transcribe a clinical encounter using Medical Mode with speaker labels and entity detection, then generate a structured SOAP note with LLM Gateway. Transcribe a clinical encounter using Medical Mode with speaker labels and entity detection, identify speakers by role, then use LLM Gateway to generate a structured SOAP note. **Products used:** [Pre-recorded STT](/pre-recorded-audio/getting-started/transcribe-an-audio-file) + [Medical Mode](/pre-recorded-audio/medical-mode) + [speaker diarization](/pre-recorded-audio/label-speakers) + [Speaker Identification](/speech-understanding/speaker-identification) + [entity detection](/speech-understanding/entity-detection) + [LLM Gateway](/llm-gateway/quickstart) **Model selection:** Uses `universal-3-5-pro` with `"domain": "medical-v1"` for purpose-built accuracy on medical terminology, drug names, and clinical language. ```python expandable import requests import time # ── Config ──────────────────────────────────────────────────── base_url = "https://api.assemblyai.com" headers = {"authorization": "YOUR_API_KEY"} audio_url = "https://assembly.ai/wildfires.mp3" # Replace with your clinical audio # ── Step 1: Transcribe with Medical Mode + speaker labels + entity detection ── data = { "audio_url": audio_url, "domain": "medical-v1", "speaker_labels": True, "entity_detection": True, } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) response.raise_for_status() transcript_id = response.json()["id"] while True: result = requests.get(f"{base_url}/v2/transcript/{transcript_id}", headers=headers).json() if result["status"] == "completed": break elif result["status"] == "error": raise RuntimeError(f"Transcription failed: {result['error']}") time.sleep(3) # ── Step 2: Identify speakers by role ── understanding_response = requests.post( "https://llm-gateway.assemblyai.com/v1/understanding", headers=headers, json={ "transcript_id": transcript_id, "speech_understanding": { "request": { "speaker_identification": { "speaker_type": "role", "known_values": ["Provider", "Patient"], } } }, }, ) understanding_response.raise_for_status() identified = understanding_response.json() # ── Step 3: Extract detected entities ── entities = result.get("entities", []) entity_summary = "\n".join( f"- {e['entity_type']}: {e['text']}" for e in entities ) # ── Step 4: Format identified transcript ── speaker_transcript = "\n".join( f"{u['speaker']}: {u['text']}" for u in identified["utterances"] ) # ── Step 5: Generate SOAP note via LLM Gateway ── llm_response = requests.post( "https://llm-gateway.assemblyai.com/v1/chat/completions", headers=headers, json={ "model": "claude-sonnet-4-6", "messages": [ { "role": "user", "content": ( "You are a medical scribe. Given the clinical encounter transcript and " "detected entities below, generate a structured SOAP note.\n\n" "Format the note with these sections:\n" "- **Subjective**: Patient's reported symptoms and history\n" "- **Objective**: Clinical observations and measurements\n" "- **Assessment**: Diagnosis or clinical impression\n" "- **Plan**: Treatment plan, prescriptions, and follow-up\n\n" f"Detected entities:\n{entity_summary}\n\n" f"Transcript:\n{speaker_transcript}" ), } ], "max_tokens": 2000, }, ) llm_response.raise_for_status() print("=== SOAP Note ===\n") print(llm_response.json()["choices"][0]["message"]["content"]) ``` ```python expandable import assemblyai as aai import requests # ── Config ──────────────────────────────────────────────────── api_key = "YOUR_API_KEY" aai.settings.api_key = api_key headers = {"authorization": api_key} audio_url = "https://assembly.ai/wildfires.mp3" # Replace with your clinical audio # ── Step 1: Transcribe with Medical Mode + speaker labels + entity detection ── config = aai.TranscriptionConfig( domain="medical-v1", speaker_labels=True, entity_detection=True, ) transcript = aai.Transcriber().transcribe(audio_url, config) if transcript.error: raise RuntimeError(f"Transcription failed: {transcript.error}") # ── Step 2: Identify speakers by role ── understanding_response = requests.post( "https://llm-gateway.assemblyai.com/v1/understanding", headers=headers, json={ "transcript_id": transcript.id, "speech_understanding": { "request": { "speaker_identification": { "speaker_type": "role", "known_values": ["Provider", "Patient"], } } }, }, ) understanding_response.raise_for_status() identified = understanding_response.json() # ── Step 3: Extract detected entities ── entities = transcript.entities or [] entity_summary = "\n".join( f"- {e.entity_type}: {e.text}" for e in entities ) # ── Step 4: Format identified transcript ── speaker_transcript = "\n".join( f"{u['speaker']}: {u['text']}" for u in identified["utterances"] ) # ── Step 5: Generate SOAP note via LLM Gateway ── llm_response = requests.post( "https://llm-gateway.assemblyai.com/v1/chat/completions", headers=headers, json={ "model": "claude-sonnet-4-6", "messages": [ { "role": "user", "content": ( "You are a medical scribe. Given the clinical encounter transcript and " "detected entities below, generate a structured SOAP note.\n\n" "Format the note with these sections:\n" "- **Subjective**: Patient's reported symptoms and history\n" "- **Objective**: Clinical observations and measurements\n" "- **Assessment**: Diagnosis or clinical impression\n" "- **Plan**: Treatment plan, prescriptions, and follow-up\n\n" f"Detected entities:\n{entity_summary}\n\n" f"Transcript:\n{speaker_transcript}" ), } ], "max_tokens": 2000, }, ) llm_response.raise_for_status() print("=== SOAP Note ===\n") print(llm_response.json()["choices"][0]["message"]["content"]) ``` ```javascript expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "YOUR_API_KEY", "Content-Type": "application/json", }; const audioUrl = "https://assembly.ai/wildfires.mp3"; // Replace with your clinical audio // Step 1: Transcribe with Medical Mode + speaker labels + entity detection let res = await fetch(`${baseUrl}/v2/transcript`, { method: "POST", headers, body: JSON.stringify({ audio_url: audioUrl, domain: "medical-v1", speaker_labels: true, entity_detection: true, }), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const { id: transcriptId } = await res.json(); let result; while (true) { res = await fetch(`${baseUrl}/v2/transcript/${transcriptId}`, { headers }); result = await res.json(); if (result.status === "completed") break; if (result.status === "error") throw new Error(`Transcription failed: ${result.error}`); await new Promise((r) => setTimeout(r, 3000)); } // Step 2: Identify speakers by role const understandingRes = await fetch( "https://llm-gateway.assemblyai.com/v1/understanding", { method: "POST", headers, body: JSON.stringify({ transcript_id: transcriptId, speech_understanding: { request: { speaker_identification: { speaker_type: "role", known_values: ["Provider", "Patient"], }, }, }, }), } ); if (!understandingRes.ok) throw new Error(`Error: ${understandingRes.status}`); const identified = await understandingRes.json(); // Step 3: Extract detected entities const entities = result.entities || []; const entitySummary = entities .map((e) => `- ${e.entity_type}: ${e.text}`) .join("\n"); // Step 4: Format identified transcript const speakerTranscript = identified.utterances .map((u) => `${u.speaker}: ${u.text}`) .join("\n"); // Step 5: Generate SOAP note via LLM Gateway res = await fetch("https://llm-gateway.assemblyai.com/v1/chat/completions", { method: "POST", headers, body: JSON.stringify({ model: "claude-sonnet-4-6", messages: [ { role: "user", content: "You are a medical scribe. Given the clinical encounter transcript and " + "detected entities below, generate a structured SOAP note.\n\n" + "Format the note with these sections:\n" + "- **Subjective**: Patient's reported symptoms and history\n" + "- **Objective**: Clinical observations and measurements\n" + "- **Assessment**: Diagnosis or clinical impression\n" + "- **Plan**: Treatment plan, prescriptions, and follow-up\n\n" + `Detected entities:\n${entitySummary}\n\n` + `Transcript:\n${speakerTranscript}`, }, ], max_tokens: 2000, }), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const llmResult = await res.json(); console.log("=== SOAP Note ===\n"); console.log(llmResult.choices[0].message.content); ``` ```javascript expandable import { AssemblyAI } from "assemblyai"; // ── Config ──────────────────────────────────────────────────── const apiKey = "YOUR_API_KEY"; const client = new AssemblyAI({ apiKey }); const headers = { authorization: apiKey, "Content-Type": "application/json", }; const audioUrl = "https://assembly.ai/wildfires.mp3"; // Replace with your clinical audio // Step 1: Transcribe with Medical Mode + speaker labels + entity detection const transcript = await client.transcripts.transcribe({ audio: audioUrl, domain: "medical-v1", speaker_labels: true, entity_detection: true, }); if (transcript.error) { throw new Error(`Transcription failed: ${transcript.error}`); } // Step 2: Identify speakers by role const understandingRes = await fetch( "https://llm-gateway.assemblyai.com/v1/understanding", { method: "POST", headers, body: JSON.stringify({ transcript_id: transcript.id, speech_understanding: { request: { speaker_identification: { speaker_type: "role", known_values: ["Provider", "Patient"], }, }, }, }), } ); if (!understandingRes.ok) throw new Error(`Error: ${understandingRes.status}`); const identified = await understandingRes.json(); // Step 3: Extract detected entities const entities = transcript.entities || []; const entitySummary = entities .map((e) => `- ${e.entity_type}: ${e.text}`) .join("\n"); // Step 4: Format identified transcript const speakerTranscript = identified.utterances .map((u) => `${u.speaker}: ${u.text}`) .join("\n"); // Step 5: Generate SOAP note via LLM Gateway const llmRes = await fetch( "https://llm-gateway.assemblyai.com/v1/chat/completions", { method: "POST", headers, body: JSON.stringify({ model: "claude-sonnet-4-6", messages: [ { role: "user", content: "You are a medical scribe. Given the clinical encounter transcript and " + "detected entities below, generate a structured SOAP note.\n\n" + "Format the note with these sections:\n" + "- **Subjective**: Patient's reported symptoms and history\n" + "- **Objective**: Clinical observations and measurements\n" + "- **Assessment**: Diagnosis or clinical impression\n" + "- **Plan**: Treatment plan, prescriptions, and follow-up\n\n" + `Detected entities:\n${entitySummary}\n\n` + `Transcript:\n${speakerTranscript}`, }, ], max_tokens: 2000, }), } ); if (!llmRes.ok) throw new Error(`Error: ${llmRes.status}`); const llmResult = await llmRes.json(); console.log("=== SOAP Note ===\n"); console.log(llmResult.choices[0].message.content); ``` ```text expandable === SOAP Note === ## Subjective Patient reports exposure to wildfire smoke over the past several days. Describes worsening cough, shortness of breath, and eye irritation. Symptoms began approximately 3 days ago coinciding with elevated air quality alerts in the region. ## Objective - AQI reading: 150 micrograms per cubic meter (10x annual average) - Particulate matter levels classified as "unhealthy" - Patient appears alert and oriented ## Assessment Acute respiratory irritation secondary to wildfire smoke exposure. Environmental exposure consistent with regional air quality emergency. ## Plan 1. Advise patient to remain indoors with windows closed 2. Recommend N95 mask for any necessary outdoor activity 3. Prescribe albuterol inhaler PRN for acute bronchospasm 4. Follow up in 1 week or sooner if symptoms worsen 5. Refer to pulmonology if symptoms persist beyond 2 weeks ``` For more on building clinical documentation apps, see the [Medical Scribe guides](/medical-scribe-best-practices). --- See the [End-to-end examples overview](/getting-started/end-to-end-examples) for all available pipelines. --- # Content repurposing URL: https://www.assemblyai.com/docs/getting-started/end-to-end-examples/content-repurposing Source: docs/getting-started/end-to-end-examples/content-repurposing.mdx Navigation: Overview > Getting started > End-to-end examples Description: Transcribe a podcast or webinar, extract key phrases, then generate a blog post draft with LLM Gateway. Transcribe a podcast or webinar, extract key phrases, then use LLM Gateway to generate a blog post draft with highlights. **Products used:** [Pre-recorded STT](/pre-recorded-audio/getting-started/transcribe-an-audio-file) + [key phrases](/speech-understanding/key-phrases) + [LLM Gateway](/llm-gateway/quickstart) **Model selection:** Uses `universal-3-5-pro` with `universal-2` fallback for multilingual content. If your content is English-only, `universal-3-5-pro` alone gives the best results. ```python expandable import requests import time # ── Config ──────────────────────────────────────────────────── base_url = "https://api.assemblyai.com" headers = {"authorization": "YOUR_API_KEY"} audio_url = "https://assembly.ai/wildfires.mp3" # ── Step 1: Transcribe with key phrases enabled ── data = { "audio_url": audio_url, "language_detection": True, "auto_highlights": True, } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) response.raise_for_status() transcript_id = response.json()["id"] while True: result = requests.get(f"{base_url}/v2/transcript/{transcript_id}", headers=headers).json() if result["status"] == "completed": break elif result["status"] == "error": raise RuntimeError(f"Transcription failed: {result['error']}") time.sleep(3) # ── Step 2: Extract top key phrases ── highlights = result.get("auto_highlights_result", {}).get("results", []) top_phrases = sorted(highlights, key=lambda x: x["rank"], reverse=True)[:10] phrases_list = ", ".join(p["text"] for p in top_phrases) # ── Step 3: Generate blog post via LLM Gateway ── llm_response = requests.post( "https://llm-gateway.assemblyai.com/v1/chat/completions", headers=headers, json={ "model": "claude-sonnet-4-6", "messages": [ { "role": "user", "content": ( "You are a content writer. Transform this transcript into an engaging " "blog post.\n\n" "Requirements:\n" "- Write a compelling title and subtitle\n" "- Break the content into 3-5 sections with headers\n" "- Weave in the key phrases naturally\n" "- Add a TL;DR at the top\n" "- End with a call-to-action\n\n" f"Key phrases: {phrases_list}\n\n" "Transcript:\n{{ transcript }}" ), } ], "transcript_id": transcript_id, "max_tokens": 3000, }, ) llm_response.raise_for_status() print("=== Blog Post Draft ===\n") print(llm_response.json()["choices"][0]["message"]["content"]) ``` ```javascript expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "YOUR_API_KEY", "Content-Type": "application/json", }; const audioUrl = "https://assembly.ai/wildfires.mp3"; // Step 1: Transcribe with key phrases enabled let res = await fetch(`${baseUrl}/v2/transcript`, { method: "POST", headers, body: JSON.stringify({ audio_url: audioUrl, language_detection: true, auto_highlights: true, }), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const { id: transcriptId } = await res.json(); let result; while (true) { res = await fetch(`${baseUrl}/v2/transcript/${transcriptId}`, { headers }); result = await res.json(); if (result.status === "completed") break; if (result.status === "error") throw new Error(`Transcription failed: ${result.error}`); await new Promise((r) => setTimeout(r, 3000)); } // Step 2: Extract top key phrases const highlights = result.auto_highlights_result?.results || []; const topPhrases = highlights .sort((a, b) => b.rank - a.rank) .slice(0, 10); const phrasesList = topPhrases.map((p) => p.text).join(", "); // Step 3: Generate blog post via LLM Gateway res = await fetch("https://llm-gateway.assemblyai.com/v1/chat/completions", { method: "POST", headers, body: JSON.stringify({ model: "claude-sonnet-4-6", messages: [ { role: "user", content: "You are a content writer. Transform this transcript into an engaging " + "blog post.\n\n" + "Requirements:\n" + "- Write a compelling title and subtitle\n" + "- Break the content into 3-5 sections with headers\n" + "- Weave in the key phrases naturally\n" + "- Add a TL;DR at the top\n" + "- End with a call-to-action\n\n" + `Key phrases: ${phrasesList}\n\n` + "Transcript:\n{{ transcript }}", }, ], transcript_id: transcriptId, max_tokens: 3000, }), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const llmResult = await res.json(); console.log("=== Blog Post Draft ===\n"); console.log(llmResult.choices[0].message.content); ``` ```text expandable === Blog Post Draft === # When the Sky Turns Orange: Understanding Wildfire Smoke and Air Quality **How Canadian wildfires are reshaping air quality across the United States** **TL;DR:** Wildfire smoke from Canada is triggering widespread air quality alerts in the US, with particulate matter levels reaching 10x normal in some cities. Here's what you need to know about the health risks and how to protect yourself. ## The smoke crosses borders Hundreds of wildfires burning across Canada have sent massive plumes of smoke southward into the United States, creating hazy skies and triggering air quality alerts from the Midwest to the Eastern Seaboard... ## Understanding particulate matter The real danger lies in fine particulate matter — microscopic particles that can penetrate deep into your lungs and even enter your bloodstream... ## Protecting your health Health experts recommend staying indoors, using air purifiers, and wearing N95 masks when outdoor exposure is unavoidable... ## Looking ahead As climate change intensifies wildfire seasons, these cross-border smoke events are likely to become more frequent... --- *Want to transcribe your own podcast or webinar? Get started with AssemblyAI's API at [assemblyai.com](https://assemblyai.com).* ``` --- See the [End-to-end examples overview](/getting-started/end-to-end-examples) for all available pipelines. --- # Real-time meeting assistant URL: https://www.assemblyai.com/docs/getting-started/end-to-end-examples/real-time-meeting-assistant Source: docs/getting-started/end-to-end-examples/real-time-meeting-assistant.mdx Navigation: Overview > Getting started > End-to-end examples Description: Stream audio from your microphone with speaker diarization and LLM Gateway for live transcription and automatic summaries. Stream audio from your microphone with speaker diarization and LLM Gateway to get live transcription and automatic summaries after each speaker turn. **Products used:** [Real-time STT](/streaming/getting-started/transcribe-streaming-audio) + [Universal-3.5 Pro](/streaming/getting-started/transcribe-streaming-audio) + [LLM Gateway](/guides/real_time_llm_gateway) **Model selection:** Uses `universal-3-5-pro` (Universal-3.5 Pro Streaming) for the lowest latency (~300ms) with the highest streaming accuracy. ```python expandable # pip install pyaudio websocket-client import pyaudio import websocket import json import threading import time from urllib.parse import urlencode # ── Config ──────────────────────────────────────────────────── YOUR_API_KEY = "YOUR_API_KEY" PROMPT = ( "Summarize this speaker turn in one sentence, then list any " "action items mentioned.\n\nTranscript: {{turn}}" ) LLM_GATEWAY_CONFIG = { "model": "claude-sonnet-4-6", "messages": [{"role": "user", "content": PROMPT}], "max_tokens": 500, } CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "format_turns": True, "min_turn_silence": 560, # Wait longer for natural meeting pauses "max_turn_silence": 2000, "llm_gateway": json.dumps(LLM_GATEWAY_CONFIG), } API_ENDPOINT = f"wss://streaming.assemblyai.com/v3/ws?{urlencode(CONNECTION_PARAMS)}" # Audio settings FRAMES_PER_BUFFER = 800 SAMPLE_RATE = 16000 stop_event = threading.Event() def on_open(ws): print("Connected — speak into your microphone. Press Ctrl+C to stop.\n") def stream_audio(): audio = pyaudio.PyAudio() stream = audio.open( input=True, frames_per_buffer=FRAMES_PER_BUFFER, channels=1, format=pyaudio.paInt16, rate=SAMPLE_RATE, ) while not stop_event.is_set(): try: data = stream.read(FRAMES_PER_BUFFER, exception_on_overflow=False) ws.send(data, websocket.ABNF.OPCODE_BINARY) except Exception: break stream.stop_stream() stream.close() audio.terminate() threading.Thread(target=stream_audio, daemon=True).start() def on_message(ws, message): data = json.loads(message) msg_type = data.get("type") if msg_type == "Turn": transcript = data.get("transcript", "") if data.get("end_of_turn") and transcript: print(f"[Turn] {transcript}\n") elif transcript: print(f"\r ... {transcript[-80:]}", end="", flush=True) elif msg_type == "LLMGatewayResponse": content = data.get("data", {}).get("choices", [{}])[0].get("message", {}).get("content", "") print(f"[Assistant] {content}\n") elif msg_type == "Termination": print(f"\nSession ended — {data.get('audio_duration_seconds', 0)}s of audio processed.") def on_error(ws, error): print(f"Error: {error}") stop_event.set() def on_close(ws, code, msg): stop_event.set() ws_app = websocket.WebSocketApp( API_ENDPOINT, header={"Authorization": YOUR_API_KEY}, on_open=on_open, on_message=on_message, on_error=on_error, on_close=on_close, ) ws_thread = threading.Thread(target=ws_app.run_forever, daemon=True) ws_thread.start() try: while ws_thread.is_alive(): time.sleep(0.1) except KeyboardInterrupt: print("\nStopping...") stop_event.set() if ws_app.sock and ws_app.sock.connected: ws_app.send(json.dumps({"type": "Terminate"})) time.sleep(2) ws_app.close() ``` ```javascript expandable // npm install ws mic const WebSocket = require("ws"); const mic = require("mic"); const querystring = require("querystring"); // ── Config ──────────────────────────────────────────────────── const YOUR_API_KEY = "YOUR_API_KEY"; const PROMPT = "Summarize this speaker turn in one sentence, then list any " + "action items mentioned.\n\nTranscript: {{turn}}"; const LLM_GATEWAY_CONFIG = { model: "claude-sonnet-4-6", messages: [{ role: "user", content: PROMPT }], max_tokens: 500, }; const CONNECTION_PARAMS = { sample_rate: 16000, speech_model: "universal-3-5-pro", format_turns: true, min_turn_silence: 560, max_turn_silence: 2000, llm_gateway: JSON.stringify(LLM_GATEWAY_CONFIG), }; const API_ENDPOINT = `wss://streaming.assemblyai.com/v3/ws?${querystring.stringify(CONNECTION_PARAMS)}`; // Connect and stream const ws = new WebSocket(API_ENDPOINT, { headers: { Authorization: YOUR_API_KEY }, }); let micInstance; ws.on("open", () => { console.log("Connected — speak into your microphone. Press Ctrl+C to stop.\n"); micInstance = mic({ rate: "16000", channels: "1", debug: false }); const micStream = micInstance.getAudioStream(); micStream.on("data", (data) => { if (ws.readyState === WebSocket.OPEN) ws.send(data); }); micInstance.start(); }); ws.on("message", (message) => { const data = JSON.parse(message); if (data.type === "Turn") { const transcript = data.transcript || ""; if (data.end_of_turn && transcript) { console.log(`[Turn] ${transcript}\n`); } else if (transcript) { process.stdout.write(`\r ... ${transcript.slice(-80)}`); } } else if (data.type === "LLMGatewayResponse") { const content = data.data?.choices?.[0]?.message?.content || ""; console.log(`[Assistant] ${content}\n`); } else if (data.type === "Termination") { console.log( `\nSession ended — ${data.audio_duration_seconds || 0}s of audio processed.` ); } }); ws.on("error", (err) => console.error(`Error: ${err}`)); ws.on("close", () => { if (micInstance) micInstance.stop(); console.log("Disconnected."); }); process.on("SIGINT", () => { console.log("\nStopping..."); if (ws.readyState === WebSocket.OPEN) { ws.send(JSON.stringify({ type: "Terminate" })); } setTimeout(() => { if (micInstance) micInstance.stop(); ws.close(); process.exit(0); }, 2000); }); ``` ```text Connected — speak into your microphone. Press Ctrl+C to stop. [Turn] So the main thing we need to decide today is whether we're going with vendor A or vendor B for the new analytics platform. [Assistant] The speaker is initiating a decision discussion about choosing between two analytics platform vendors. Action items: None yet — decision pending. [Turn] I think vendor A has better pricing but vendor B has the integrations we need. Can someone pull the comparison spreadsheet by Friday? [Assistant] The speaker compared vendor pricing vs. integrations and requested a comparison document. Action items: - Pull the vendor comparison spreadsheet by Friday ``` --- See the [End-to-end examples overview](/getting-started/end-to-end-examples) for all available pipelines. --- # Real-time live captioner URL: https://www.assemblyai.com/docs/getting-started/end-to-end-examples/real-time-live-captioner Source: docs/getting-started/end-to-end-examples/real-time-live-captioner.mdx Navigation: Overview > Getting started > End-to-end examples Description: Stream audio from your microphone with keyterms prompting for domain-specific accuracy, ideal for live events and accessibility. Stream audio from your microphone with keyterms prompting for domain-specific accuracy, ideal for live events, accessibility, and broadcast captioning. **Products used:** [Real-time STT](/streaming/getting-started/transcribe-streaming-audio) + [Universal-3.5 Pro](/streaming/getting-started/transcribe-streaming-audio) + [keyterms prompting](/streaming/prompting-and-keyterms) **Model selection:** Uses `universal-3-5-pro` for sub-300ms latency with `format_turns` enabled for clean, readable captions. ```python expandable # pip install pyaudio websocket-client import pyaudio import websocket import json import threading import time from urllib.parse import urlencode # ── Config ──────────────────────────────────────────────────── YOUR_API_KEY = "YOUR_API_KEY" # Add domain-specific terms to boost recognition accuracy KEYTERMS = ["AssemblyAI", "Universal-3.5 Pro", "LLM Gateway", "speech-to-text"] CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "format_turns": True, "keyterms_prompt": KEYTERMS, } API_ENDPOINT = ( f"wss://streaming.assemblyai.com/v3/ws?{urlencode(CONNECTION_PARAMS, doseq=True)}" ) # Audio settings FRAMES_PER_BUFFER = 800 SAMPLE_RATE = 16000 stop_event = threading.Event() caption_count = 0 def on_open(ws): print(f"Live captioning started — keyterms: {', '.join(KEYTERMS)}") print("Speak into your microphone. Press Ctrl+C to stop.\n") print("-" * 60) def stream_audio(): audio = pyaudio.PyAudio() stream = audio.open( input=True, frames_per_buffer=FRAMES_PER_BUFFER, channels=1, format=pyaudio.paInt16, rate=SAMPLE_RATE, ) while not stop_event.is_set(): try: data = stream.read(FRAMES_PER_BUFFER, exception_on_overflow=False) ws.send(data, websocket.ABNF.OPCODE_BINARY) except Exception: break stream.stop_stream() stream.close() audio.terminate() threading.Thread(target=stream_audio, daemon=True).start() def on_message(ws, message): global caption_count data = json.loads(message) if data.get("type") == "Turn": transcript = data.get("transcript", "") if data.get("end_of_turn") and transcript: caption_count += 1 print(f"\r[{caption_count:03d}] {transcript}") elif transcript: # Show partial (live) caption print(f"\r >> {transcript[-70:]}", end="", flush=True) elif data.get("type") == "Termination": duration = data.get("audio_duration_seconds", 0) print(f"\n{'=' * 60}") print(f"Session ended — {caption_count} captions, {duration}s of audio") def on_error(ws, error): print(f"\nError: {error}") stop_event.set() def on_close(ws, code, msg): stop_event.set() ws_app = websocket.WebSocketApp( API_ENDPOINT, header={"Authorization": YOUR_API_KEY}, on_open=on_open, on_message=on_message, on_error=on_error, on_close=on_close, ) ws_thread = threading.Thread(target=ws_app.run_forever, daemon=True) ws_thread.start() try: while ws_thread.is_alive(): time.sleep(0.1) except KeyboardInterrupt: print("\n\nStopping...") stop_event.set() if ws_app.sock and ws_app.sock.connected: ws_app.send(json.dumps({"type": "Terminate"})) time.sleep(2) ws_app.close() ``` ```javascript expandable // npm install ws mic const WebSocket = require("ws"); const mic = require("mic"); const querystring = require("querystring"); // ── Config ──────────────────────────────────────────────────── const YOUR_API_KEY = "YOUR_API_KEY"; // Add domain-specific terms to boost recognition accuracy const KEYTERMS = ["AssemblyAI", "Universal-3.5 Pro", "LLM Gateway", "speech-to-text"]; const CONNECTION_PARAMS = { sample_rate: 16000, speech_model: "universal-3-5-pro", format_turns: true, keyterms_prompt: KEYTERMS, }; const API_ENDPOINT = `wss://streaming.assemblyai.com/v3/ws?${querystring.stringify( CONNECTION_PARAMS )}`; let micInstance; let captionCount = 0; const ws = new WebSocket(API_ENDPOINT, { headers: { Authorization: YOUR_API_KEY }, }); ws.on("open", () => { console.log(`Live captioning started — keyterms: ${KEYTERMS.join(", ")}`); console.log("Speak into your microphone. Press Ctrl+C to stop.\n"); console.log("-".repeat(60)); micInstance = mic({ rate: "16000", channels: "1", debug: false }); const micStream = micInstance.getAudioStream(); micStream.on("data", (data) => { if (ws.readyState === WebSocket.OPEN) ws.send(data); }); micInstance.start(); }); ws.on("message", (message) => { const data = JSON.parse(message); if (data.type === "Turn") { const transcript = data.transcript || ""; if (data.end_of_turn && transcript) { captionCount++; process.stdout.write( `\r${String(captionCount).padStart(3, "0")} ${transcript}\n` ); } else if (transcript) { process.stdout.write(`\r >> ${transcript.slice(-70)}`); } } else if (data.type === "Termination") { console.log(`\n${"=".repeat(60)}`); console.log( `Session ended — ${captionCount} captions, ${data.audio_duration_seconds || 0}s of audio` ); } }); ws.on("error", (err) => console.error(`Error: ${err}`)); ws.on("close", () => { if (micInstance) micInstance.stop(); }); process.on("SIGINT", () => { console.log("\n\nStopping..."); if (ws.readyState === WebSocket.OPEN) { ws.send(JSON.stringify({ type: "Terminate" })); } setTimeout(() => { if (micInstance) micInstance.stop(); ws.close(); process.exit(0); }, 2000); }); ``` ```text Live captioning started — keyterms: AssemblyAI, Universal-3.5 Pro, LLM Gateway, speech-to-text Speak into your microphone. Press Ctrl+C to stop. ------------------------------------------------------------ [001] Welcome everyone to today's demo of AssemblyAI's speech-to-text platform. [002] We'll be showing you how Universal-3.5 Pro handles real-time transcription. [003] The LLM Gateway integration lets you add AI analysis on top of your transcripts without switching providers. ============================================================ Session ended — 3 captions, 24s of audio ``` --- See the [End-to-end examples overview](/getting-started/end-to-end-examples) for all available pipelines. --- # Best Practices for building Meeting Notetakers URL: https://www.assemblyai.com/docs/meeting-notetaker-best-practices Source: docs/meeting-notetaker-best-practices.mdx Navigation: Overview > Use cases & integrations > Use case guides Description: Complete guide for building meeting notetakers with AssemblyAI ## Introduction Building a robust meeting notetaker requires careful consideration of accuracy, latency, speaker identification, and real-time capabilities. This guide addresses common questions and provides practical solutions for both post-call and live meeting transcription scenarios. ## Why AssemblyAI for Meeting Notetakers? AssemblyAI stands out as the premier choice for meeting notetakers with several key advantages: ### Industry-Leading Accuracy with Pre-recorded Audio - **93.3%+ transcription accuracy** ensures reliable meeting documentation - **2.9% speaker diarization error rate** for precise "who said what" attribution - **Speech Understanding** integration for intelligent post-processing and insights - **Keyterms prompt** allows providing meeting context to improve accuracy of transcription ### Streaming with Universal-3.5 Pro As meeting notetakers evolve toward real-time capabilities, AssemblyAI's Universal-3.5 Pro Streaming model (`universal-3-5-pro`) offers significant benefits: - **Speaker diarization** available for both pre-recorded and streaming transcription - **Ultra-low latency (~300ms)** enables live transcription without delays - **Format turns** feature provides structured, readable output in real-time - **Keyterms prompt** allows providing meeting context to improve accuracy of transcription ### End-to-End Voice AI Platform Unlike fragmented solutions, AssemblyAI provides a unified API for: - Transcription with speaker diarization - Automatic language detection and code switching - Boosting accuracy via meeting context with keyterms prompt - Speech Understanding tasks like speaker identification, translation, and transcript styling - Post-processing workflows with custom prompting - from summarization to completely custom workflows - Real-time and batch processing of pre-recorded audio in a single platform ## When Should I Use Pre-recorded vs Streaming for Meeting Notetakers? Understanding when to use pre-recorded versus streaming speech-to-text is critical for building the right meeting notetaker. ### Pre-recorded Speech-to-text **Post-call analysis** - Meeting already happened, you have the full recording - **Highest accuracy needed** - Pre-recorded models have higher accuracy (93.3%+) - **Speaker diarization is critical** - Pre-recorded has 2.9% speaker error rate - **Broad language support** - Need any of 99+ languages - **Advanced features required** - Summarization, sentiment analysis, entity detection, PII redaction, speaker identification - **Batch processing** - Processing multiple recordings at once - **Quality over speed** - Can wait seconds/minutes for perfect results **Best for:** Zoom/Teams/Meet recording uploads, compliance, documentation, post-call summaries, searchable archives ### Streaming Speech-to-text **Live meetings** - Transcribing as the meeting happens You should use streaming when you need to display a live transcript of text to users as they are speaking. With Universal-3.5 Pro Streaming, accuracy is closer to pre-recorded, but pre-recorded will always be the most accurate option. - **Real-time captions** - Displaying subtitles/captions to participants during calls - **Immediate feedback** - Need transcription within ~300ms - **Interactive features** - Live note-taking, real-time keyword detection, action item alerts - **No recording available** - Processing live audio only **Best for:** Live captions, real-time note-taking apps, accessibility features, live keyword alerts **Streaming is billed per session** Streaming is billed on the total duration that your WebSocket connection stays open, not on the amount of audio you send. For long-running meetings, make sure to terminate sessions when the meeting ends to avoid being billed for idle time. See [Billing and pricing](/billing-and-pricing) for details. ### Hybrid Approach (Recommended) Many successful meeting notetakers use **both** pre-recorded and streaming speech-to-text: 1. **Streaming during the call** - Provide live captions and real-time notes to participants 2. **Pre-recorded after the call** - Generate high-quality transcript with speaker labels, summary, and insights This gives users immediate value during meetings while providing comprehensive documentation afterward. **Example workflow:** - User joins meeting → Start streaming for live captions - Meeting ends → Upload recording to pre-recorded API for final transcript with speaker names - Generate meeting summary, action items, and searchable archive from pre-recorded transcript ## What Languages and Features for a Meeting Notetaker? ### Pre-Recorded Meetings For post-call analysis, AssemblyAI supports: **Languages**: - 99 languages supported - Automatic Language Detection to route to the most spoken language - Code Switching to preserve changes in speech between languages **Core Features**: - Speaker diarization (1-10 speakers by default, expandable to any min/max) - Multichannel audio support (each channel = one speaker) - Automatic formatting, punctuation, and capitalization - Keyterms prompting for boosting domain-specific terms **Speech Understanding Models**: - Summarization for meeting recaps - Sentiment analysis for meeting tone assessment - Entity detection for extracting key information - Speaker identification to map generic labels to actual names/roles - Translation between 86 languages ### Real-Time Streaming For live meeting transcription: **Languages**: - English-only model (default) - Multilingual model supporting English, Spanish, French, German, Portuguese, and Italian ### Streaming (Universal-3.5 Pro Streaming) - Speaker diarization for identifying who is speaking - Partial and final transcripts for responsive UI - Format turns for structured, readable output - Keyterms prompt for contextual accuracy See the [Universal-3.5 Pro Streaming documentation](/streaming/getting-started/transcribe-streaming-audio) for full details. ## How Can I Get Started Building a Post-Call Meeting Notetaker? Here's a complete example implementing pre-recorded transcription with all essential features: ```python expandable import assemblyai as aai import asyncio from typing import Dict, List from assemblyai.types import ( SpeakerOptions, LanguageDetectionOptions, PIIRedactionPolicy, PIISubstitutionPolicy, ) # Configure API key aai.settings.api_key = "your_api_key_here" async def transcribe_meeting_async(audio_source: str) -> Dict: """ Asynchronously transcribe a meeting recording with full features Args: audio_source: Either a local file path or publicly accessible URL """ # Configure comprehensive meeting analysis config = aai.TranscriptionConfig( # Speaker diarization speaker_labels=True, speakers_expected=None, # Use if you know exact number from Zoom/Meet/Teams speaker_options=SpeakerOptions( min_speakers_expected=2, max_speakers_expected=10 # Set a bit higher than expected; too high can cause over-splitting ), multichannel=False, # Set to True if audio has separate channel per speaker # Language detection language_detection=True, # Auto-detect the most used language language_detection_options=LanguageDetectionOptions( code_switching=True, # Preserve language switches code_switching_confidence_threshold=0.5, ), # Punctuation and formatting punctuate=True, format_text=True, # Boost accuracy of meeting-specific vocabulary keyterms_prompt=["quarterly", "KPI", "roadmap", "deliverables"], # Speech Understanding - commonly used models summarization=True, sentiment_analysis=True, entity_detection=True, redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.organization, PIIRedactionPolicy.occupation, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True ) # Create transcriber transcriber = aai.Transcriber() try: # Submit transcription job transcript = await asyncio.to_thread( transcriber.transcribe, audio_source, config=config ) # Check status if transcript.status == aai.TranscriptStatus.error: raise Exception(f"Transcription failed: {transcript.error}") # Process speaker-labeled utterances print("\n=== SPEAKER-LABELED TRANSCRIPT ===\n") for utterance in transcript.utterances: # Format timestamp start_time = utterance.start / 1000 # Convert to seconds end_time = utterance.end / 1000 # Print formatted utterance print(f"[{start_time:.1f}s - {end_time:.1f}s] Speaker {utterance.speaker}:") print(f" {utterance.text}") print(f" Confidence: {utterance.confidence:.2%}\n") # Print summary data print("\n=== MEETING SUMMARY ===\n") print({ "id": transcript.id, "status": transcript.status, "duration": transcript.audio_duration, "speaker_count": len(set(u.speaker for u in transcript.utterances)), "word_count": len(transcript.words) if transcript.words else 0, "detected_language": transcript.language_code if hasattr(transcript, 'language_code') else None, "summary": transcript.summary, }) return { "transcript": transcript, "utterances": transcript.utterances, "summary": transcript.summary, } except Exception as e: print(f"Error during transcription: {e}") raise async def main(): """ Example usage with error handling """ # Use either local file OR URL (not both) audio_source = "https://assembly.ai/wildfires.mp3" # Or "path/to/recording.mp3" try: result = await transcribe_meeting_async(audio_source) # Additional processing print(f"\nTotal speakers identified: {len(set(u.speaker for u in result['utterances']))}") print(f"Meeting duration: {result['transcript'].audio_duration} seconds") except Exception as e: print(f"Failed to process meeting: {e}") if __name__ == "__main__": asyncio.run(main()) ``` ## How Can I Get Started Building a During-Call Live Meeting Notetaker? Here's a complete example for real-time streaming transcription with meeting-optimized settings: ```python expandable # pip install pyaudio websocket-client import pyaudio import websocket import json import threading import time from urllib.parse import urlencode from datetime import datetime # --- Configuration --- YOUR_API_KEY = "your_api_key" # Keyterms to improve recognition accuracy KEYTERMS = [ "Alice Johnson", "Bob Smith", "Carol Davis", "quarterly review", "action items", "follow up", "deadline", "budget" ] # MEETING NOTETAKER CONFIGURATION (different from voice agents!) CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "format_turns": True, # ALWAYS TRUE for meetings - users need readable text # Meeting-optimized turn detection (wait longer than voice agents) # universal-3-5-pro defaults: min_turn_silence=100ms, max_turn_silence=1000ms "min_turn_silence": 560, # Wait longer for natural pauses (voice agents use ~100ms) "max_turn_silence": 2000, # Allow thinking pauses # Keyterms for accuracy - pass each term as a separate query parameter "keyterms_prompt": KEYTERMS, } API_ENDPOINT_BASE_URL = "wss://streaming.assemblyai.com/v3/ws" API_ENDPOINT = f"{API_ENDPOINT_BASE_URL}?{urlencode(CONNECTION_PARAMS, doseq=True)}" # Audio Configuration FRAMES_PER_BUFFER = 800 # 50ms of audio SAMPLE_RATE = CONNECTION_PARAMS["sample_rate"] CHANNELS = 1 FORMAT = pyaudio.paInt16 # Global variables audio = None stream = None ws_app = None audio_thread = None stop_event = threading.Event() transcript_buffer = [] def on_open(ws): """Called when the WebSocket connection is established.""" print("=" * 80) print(f"[{datetime.now().strftime('%H:%M:%S')}] Meeting transcription started") print(f"Connected to: {API_ENDPOINT_BASE_URL}") print(f"Keyterms configured: {', '.join(KEYTERMS)}") print("=" * 80) print("\nSpeak into your microphone. Press Ctrl+C to stop.\n") def stream_audio(): """Stream audio from microphone to WebSocket""" global stream while not stop_event.is_set(): try: audio_data = stream.read(FRAMES_PER_BUFFER, exception_on_overflow=False) ws.send(audio_data, websocket.ABNF.OPCODE_BINARY) except Exception as e: if not stop_event.is_set(): print(f"Error streaming audio: {e}") break global audio_thread audio_thread = threading.Thread(target=stream_audio) audio_thread.daemon = True audio_thread.start() def on_message(ws, message): """Handle incoming messages from AssemblyAI""" try: data = json.loads(message) msg_type = data.get("type") # Uncomment to see full JSON for debugging: # print("=" * 80) # print(json.dumps(data, indent=2, ensure_ascii=False)) # print("=" * 80) # print() if msg_type == "Begin": session_id = data.get("id", "N/A") print(f"[SESSION] Started - ID: {session_id}\n") elif msg_type == "Turn": end_of_turn = data.get("end_of_turn", False) transcript = data.get("transcript", "") turn_order = data.get("turn_order", 0) end_of_turn_confidence = data.get("end_of_turn_confidence", 0.0) # FOR MEETING NOTETAKERS: Show partials for responsive UI if not end_of_turn and transcript: print(f"\r[LIVE] {transcript}", end="", flush=True) # FOR MEETING NOTETAKERS: Use formatted finals for readable display # (Unlike voice agents which should use utterance for speed) if end_of_turn and transcript: timestamp = datetime.now().strftime('%H:%M:%S') print(f"\n[{timestamp}] {transcript}") print(f" Turn: {turn_order} | Confidence: {end_of_turn_confidence:.2%}") # Detect action items transcript_lower = transcript.lower() if any(term in transcript_lower for term in ["action item", "follow up", "deadline", "assigned to", "todo"]): print(" ⚠️ ACTION ITEM DETECTED!") # Store final transcript transcript_buffer.append({ "timestamp": timestamp, "text": transcript, "turn_order": turn_order, "confidence": end_of_turn_confidence, "type": "final" }) print() elif msg_type == "Termination": audio_duration = data.get("audio_duration_seconds", 0) print(f"\n[SESSION] Terminated - Duration: {audio_duration}s") save_transcript() elif msg_type == "Error": error_msg = data.get("error", "Unknown error") print(f"\n[ERROR] {error_msg}") except json.JSONDecodeError as e: print(f"Error decoding message: {e}") except Exception as e: print(f"Error handling message: {e}") def on_error(ws, error): """Called when a WebSocket error occurs.""" print(f"\n[WEBSOCKET ERROR] {error}") stop_event.set() def on_close(ws, close_status_code, close_msg): """Called when the WebSocket connection is closed.""" print(f"\n[WEBSOCKET] Disconnected - Status: {close_status_code}, Message: {close_msg}") global stream, audio stop_event.set() # Clean up audio stream if stream: if stream.is_active(): stream.stop_stream() stream.close() stream = None if audio: audio.terminate() audio = None if audio_thread and audio_thread.is_alive(): audio_thread.join(timeout=1.0) def save_transcript(): """Save the transcript to a file""" if not transcript_buffer: print("No transcript to save.") return filename = f"meeting_transcript_{datetime.now().strftime('%Y%m%d_%H%M%S')}.txt" with open(filename, "w", encoding="utf-8") as f: f.write("Meeting Transcript\n") f.write(f"Date: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n") f.write(f"Keyterms: {', '.join(KEYTERMS)}\n") f.write("=" * 80 + "\n\n") for entry in transcript_buffer: f.write(f"[{entry['timestamp']}] {entry['text']}\n") f.write(f"Confidence: {entry['confidence']:.2%}\n\n") print(f"Transcript saved to: {filename}") def run(): """Main function to run the streaming transcription""" global audio, stream, ws_app # Initialize PyAudio audio = pyaudio.PyAudio() # Open microphone stream try: stream = audio.open( input=True, frames_per_buffer=FRAMES_PER_BUFFER, channels=CHANNELS, format=FORMAT, rate=SAMPLE_RATE, ) print("Microphone stream opened successfully.") except Exception as e: print(f"Error opening microphone stream: {e}") if audio: audio.terminate() return # Create WebSocketApp ws_app = websocket.WebSocketApp( API_ENDPOINT, header={"Authorization": YOUR_API_KEY}, on_open=on_open, on_message=on_message, on_error=on_error, on_close=on_close, ) # Run WebSocketApp in a separate thread ws_thread = threading.Thread(target=ws_app.run_forever) ws_thread.daemon = True ws_thread.start() try: # Keep main thread alive until interrupted while ws_thread.is_alive(): time.sleep(0.1) except KeyboardInterrupt: print("\n\nCtrl+C received. Stopping transcription...") stop_event.set() # Send termination message to the server if ws_app and ws_app.sock and ws_app.sock.connected: try: terminate_message = {"type": "Terminate"} ws_app.send(json.dumps(terminate_message)) time.sleep(1) except Exception as e: print(f"Error sending termination message: {e}") if ws_app: ws_app.close() ws_thread.join(timeout=2.0) finally: # Final cleanup if stream and stream.is_active(): stream.stop_stream() if stream: stream.close() if audio: audio.terminate() print("Cleanup complete. Exiting.") if __name__ == "__main__": run() ``` These settings wait longer before ending turns to accommodate natural conversation pauses and ensure readable formatted text for display. You can [tweak these settings](/streaming/getting-started/transcribe-streaming-audio) to get the best results for your notetaker. ## How Do I Handle Multichannel Meeting Audio? Many meeting platforms (Zoom, Teams, Google Meet) can record each participant on separate audio channels. This dramatically improves speaker identification accuracy. ### For Pre-recorded Meetings ```python config = aai.TranscriptionConfig( multichannel=True, # Enable when each speaker is on different channel speaker_labels=False, # Disable - channels already separate speakers # Other settings... ) transcriber = aai.Transcriber() transcript = transcriber.transcribe(audio_file, config=config) # Access per-channel transcripts for channel, channel_transcript in enumerate(transcript.channels): print(f"\n=== Channel {channel} ===") print(channel_transcript.text) ``` **When to use multichannel:** - Zoom local recordings with "Record separate audio file for each participant" enabled - Professional podcast recordings with individual microphones - Conference systems with dedicated channels per participant - Phone calls with caller and callee on separate channels **Benefits:** - **Perfect speaker separation** - No diarization errors - **No speaker confusion or overlap issues** - **Faster processing time** - Diarization not needed - **Higher accuracy** - Model processes clean single-speaker audio **How to enable in meeting platforms:** - **Zoom**: Settings → Recording → Advanced → "Record a separate audio file for each participant" - **Teams**: Requires third-party recording solutions like [Recall.ai](https://www.recall.ai/) - **Google Meet**: Requires third-party recording solutions like [Recall.ai](https://www.recall.ai/) ### For Streaming Meetings For real-time multichannel audio, create separate streaming sessions per channel: ```python expandable import asyncio import websockets class ChannelTranscriber: def __init__(self, channel_id: int, speaker_name: str): self.channel_id = channel_id self.speaker_name = speaker_name self.connection_params = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "format_turns": True, } async def transcribe_channel(self, audio_stream): """Transcribe a single audio channel""" url = f"wss://streaming.assemblyai.com/v3/ws?{urlencode(self.connection_params)}" # If you're using `websockets` version 13.0 or later, use `additional_headers` parameter. For older versions (< 13.0), use `extra_headers` instead. async with websockets.connect(url, additional_headers={"Authorization": API_KEY}) as ws: # Send audio from this channel only async for audio_chunk in audio_stream: await ws.send(audio_chunk) # Receive transcripts async for message in ws: data = json.loads(message) if data.get("type") == "Turn" and data.get("end_of_turn"): print(f"{self.speaker_name}: {data['transcript']}") # Create transcriber for each channel async def transcribe_multichannel_meeting(channel_audio_streams): transcribers = [ ChannelTranscriber(0, "Alice"), ChannelTranscriber(1, "Bob"), ] # Run all channels concurrently await asyncio.gather(*[ t.transcribe_channel(stream) for t, stream in zip(transcribers, channel_audio_streams) ]) ``` See our [multichannel streaming guide](/streaming/label-speakers-and-separate-channels#multichannel-streaming-audio) for complete implementation details. ## How Should I Handle Pre-recorded Transcription in Production? Choose the right approach based on your application's needs: ### Option 1: Simple Blocking Call ```python # Simple blocking call transcript = await asyncio.to_thread(transcriber.transcribe, audio_url, config=config) ``` **Pros:** - Simple, straightforward code - Good for low volume applications - Easy to understand and debug **Cons:** - Ties up resources while waiting - Not suitable for high volume - Cannot process multiple files simultaneously **Best for:** Personal projects, prototypes, low-traffic applications ### Option 2: Webhook Callbacks (Production Recommended) ```python config = aai.TranscriptionConfig( webhook_url="https://your-app.com/webhooks/assemblyai", webhook_auth_header_name="X-Webhook-Secret", webhook_auth_header_value="your_secret_here", speaker_labels=True, summarization=True, # ... other config ) # Submit job and return immediately (non-blocking) transcript = transcriber.submit(audio_url, config=config) print(f"Job submitted: {transcript.id}") # Your app can continue processing other requests # Your webhook receives results when ready (typically 15-30% of audio duration) ``` **Webhook handler example:** ```python expandable from flask import Flask, request, jsonify app = Flask(__name__) @app.route("/webhooks/assemblyai", methods=["POST"]) def assemblyai_webhook(): # Verify webhook authenticity if request.headers.get("X-Webhook-Secret") != "your_secret_here": return jsonify({"error": "Unauthorized"}), 401 import requests as http_requests data = request.json transcript_id = data["transcript_id"] status = data["status"] if status == "completed": # Fetch the full transcript (webhook only sends transcript_id and status) transcript = http_requests.get( f"https://api.assemblyai.com/v2/transcript/{transcript_id}", headers={"authorization": "your_api_key"} ).json() process_completed_meeting(transcript) elif status == "error": log_transcription_error(transcript_id) return jsonify({"received": True}), 200 def process_completed_meeting(transcript): """Process completed meeting transcript""" utterances = transcript["utterances"] summary = transcript["summary"] # Store in database save_to_database(transcript) # Notify user send_notification(transcript["id"]) ``` **Pros:** - Non-blocking - submit and forget - Scales to high volume - Process multiple files in parallel - Automatic retry on failures - Get notified when complete **Best for:** Production apps, user-uploaded recordings, batch processing, SaaS products ### Option 3: Polling (Custom Workflows) ```python # Submit job transcript = transcriber.submit(audio_url, config=config) print(f"Submitted: {transcript.id}") # Poll for completion with progress tracking while transcript.status not in [aai.TranscriptStatus.completed, aai.TranscriptStatus.error]: await asyncio.sleep(5) transcript = transcriber.get_transcript(transcript.id) # Optional: Show progress print(f"Status: {transcript.status}...") if transcript.status == aai.TranscriptStatus.completed: process_transcript(transcript) else: print(f"Error: {transcript.error}") ``` **Pros:** - Full control over retry logic - Can show progress to users - Good for background jobs - Works without webhook infrastructure **Cons:** - Must implement your own polling loop - Ties up resources while polling - More complex than webhooks **Best for:** Background job processors, CLIs with progress bars, custom retry logic ### Comparison Table | Method | Blocking | Scalability | Complexity | Best For | | ----------- | -------- | ----------- | ---------- | ---------------------------- | | Blocking | Yes | Low | Low | Prototypes, low volume | | Webhooks | No | High | Medium | Production, high volume | | Polling | Partial | Medium | Medium | Background jobs, progress UI | ### Scaling Considerations - **HTTP rate limit:** 20,000 requests per 5-minute window, counted across submissions (POST) and polling (GET) combined - **Exceeding the limit:** returns a `403` response - **Parallel transcriptions (rate limit):** 200+ for paid accounts (queued beyond that) - **Ramp up gradually:** start at 10-50 parallel requests, double incrementally - **Avoid the rate limit:** use [webhooks](/pre-recorded-audio/webhooks) or jittered, widened polling — see [Polling without exceeding the rate limit](/pre-recorded-audio/guides/bulk-transcription-and-load-tests-at-scale#polling-without-exceeding-the-rate-limit) - **Contact Sales** before large-scale rollouts ## How Do I Identify Speakers in My Recording? Speaker diarization tells you **when** speakers change ("Speaker A", "Speaker B"), but **Speaker Identification** tells you **who** they are by name or role. ### Why Use Speaker Identification? **Instead of:** ``` Speaker A: Let's review the Q3 numbers. Speaker B: Revenue was up 15% this quarter. Speaker A: Excellent work on that launch. ``` **You get:** ``` Sarah Chen: Let's review the Q3 numbers. Michael Rodriguez: Revenue was up 15% this quarter. Sarah Chen: Excellent work on that launch. ``` ### How It Works Speaker Identification uses AssemblyAI's Speech Understanding API to map generic speaker labels to actual names or roles that you provide: ```python expandable import assemblyai as aai aai.settings.api_key = "your_api_key" # Step 1: Transcribe with speaker diarization config = aai.TranscriptionConfig( speaker_labels=True, # Must enable speaker diarization first speech_understanding={ "request": { "speaker_identification": { "speaker_type": "name", # or "role" "known_values": ["Sarah Chen", "Michael Rodriguez", "Alex Kim"] } } } ) transcriber = aai.Transcriber() transcript = transcriber.transcribe("meeting_recording.mp3", config=config) # Access results with identified speakers for utterance in transcript.utterances: print(f"{utterance.speaker}: {utterance.text}") ``` ### Identifying by Role Instead of Name For customer service, sales calls, or scenarios where you don't know names: ```python config = aai.TranscriptionConfig( speaker_labels=True, speech_understanding={ "request": { "speaker_identification": { "speaker_type": "role", "known_values": ["Agent", "Customer"] # or ["Interviewer", "Interviewee"] } } } ) ``` **Common role combinations:** - `["Agent", "Customer"]` - Customer service calls - `["Support", "Customer"]` - Technical support - `["Interviewer", "Interviewee"]` - Interviews - `["Host", "Guest"]` - Podcasts - `["Doctor", "Patient"]` - Medical consultations (with HIPAA compliance) ### How to Get Speaker Names **For platform recordings:** 1. **Zoom**: Extract participant names from Zoom API or meeting JSON 2. **Teams**: Get attendees from Microsoft Graph API 3. **Google Meet**: Use Google Calendar API to get participants **Example with Zoom:** ```python # Get participant names from Zoom meeting zoom_participants = get_zoom_meeting_participants(meeting_id) speaker_names = [p["name"] for p in zoom_participants] # Use in speaker identification config = aai.TranscriptionConfig( speaker_labels=True, speakers_expected=len(speaker_names), # Exact number of speakers to detect speech_understanding={ "request": { "speaker_identification": { "speaker_type": "name", "known_values": speaker_names } } } ) ``` ### How Speaker Identification Works **Speaker Identification Requirements:** 1. **Speaker diarization must be enabled** - Cannot identify speakers without diarization first 2. **Requires sufficient audio per speaker** - Each speaker needs enough speech for accurate matching 3. **Works best with distinct voices** - Similar voices may be confused 4. **Post-processing step** - Adds additional processing time after transcription **Accuracy depends on:** - Audio quality (clear, minimal background noise) - Voice distinctiveness (different genders, accents, tones) - Amount of speech per speaker (more = better) - Number of speakers (fewer = more accurate) ### Alternative: Add Identification Later You can add speaker identification to an existing transcript by posting to the Speech Understanding API with the `transcript_id`. This is useful when you get speaker names after the transcription completes, or when building iterative workflows where users confirm speaker identities. ```python expandable import requests # First, transcribe with speaker diarization transcript = transcriber.transcribe(audio_url, config=aai.TranscriptionConfig(speaker_labels=True)) # Later, add speaker identification using the transcript ID understanding_body = { "transcript_id": transcript.id, "speech_understanding": { "request": { "speaker_identification": { "speaker_type": "name", "known_values": ["Sarah Chen", "Michael Rodriguez"] } } } } result = requests.post( "https://llm-gateway.assemblyai.com/v1/understanding", headers={"Authorization": aai.settings.api_key}, json=understanding_body ).json() # Access identified speakers from the response for utterance in result["utterances"]: print(f"{utterance['speaker']}: {utterance['text']}") ``` This approach is useful when: - You get speaker names after the transcription completes - You want to try different name mappings - Building iterative workflows where users confirm speaker identities For complete API details, see our [Speaker Identification documentation](/speech-understanding/speaker-identification). ## How Do I Translate Between Languages in Meetings? AssemblyAI supports translation between 86 languages, enabling you to transcribe meetings in one language and translate to another. ### When to Use Translation **Common use cases:** - Transcribe Spanish meeting → Translate to English for documentation - Transcribe multilingual meeting → Translate all to common language - Create translated meeting notes for international teams - Provide translated summaries for stakeholders ### Basic Translation Translation is a Speech Understanding feature. You enable it via the `speech_understanding` parameter with `target_languages`: ```python expandable import requests import time base_url = "https://api.assemblyai.com" headers = {"authorization": "YOUR_API_KEY"} # Configure transcription with translation data = { "audio_url": "https://assembly.ai/wildfires.mp3", "speech_models": ["universal-3-5-pro", "universal-2"], "language_detection": True, "speaker_labels": True, "speech_understanding": { "request": { "translation": { "target_languages": ["es", "de"], "formal": True } } } } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) transcript_id = response.json()["id"] polling_endpoint = base_url + f"/v2/transcript/{transcript_id}" while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) print("--- Original Transcript ---") print(transcript["text"][:200] + "...") print("\n--- Translations ---") for language_code, translated_text in transcript["translated_texts"].items(): print(f"{language_code.upper()}:") print(translated_text[:200] + "...") ``` ### Translation with Speaker Labels For meetings where you need per-utterance translations with speaker attribution: ```python data = { "audio_url": audio_url, "speech_models": ["universal-3-5-pro", "universal-2"], "speaker_labels": True, "speech_understanding": { "request": { "translation": { "target_languages": ["es"], "match_original_utterance": True, "formal": True } } } } for utterance in transcript["utterances"]: print(f"Speaker {utterance['speaker']}:") print(f" Original: {utterance['text'][:100]}...") print(f" Spanish: {utterance['translated_texts']['es'][:100]}...") ``` ### Supported Language Pairs AssemblyAI supports translation between **86 languages**, including: **Popular combinations:** - Spanish ↔ English - French ↔ English - German ↔ English - Mandarin ↔ English - Japanese ↔ English - Portuguese ↔ English - And all combinations between supported languages ### Translation Response Format The response includes `translated_texts` as a dictionary keyed by language code: ```python { "text": "Original transcript in source language", "translated_texts": { "es": "Translated transcript in Spanish", "de": "Translated transcript in German" }, "utterances": [ { "speaker": "A", "text": "Hello, how are you?", "translated_texts": { "es": "Hola, ¿cómo estás?" }, "start": 0, "end": 1500 } ] } ``` For complete language support and translation details, see our [Translation documentation](/speech-understanding/translation). ## What Workflows Can I Build for My AI Meeting Notetaker? Use these Speech Understanding and Guardrails features to transform raw transcripts into actionable insights. ### Summarization `summarization: true` **What it does:** Generates an abstractive recap of the conversation (not verbatim). **Output:** `summary` string (bullets/paragraph format). **Great for:** Meeting notes, call recaps, executive summaries. **Notes:** Condenses and rephrases; minor details may be omitted by design. **Example:** ```python config = aai.TranscriptionConfig( summarization=True, summary_type="bullets", # or "bullets_verbose", "gist", "headline", "paragraph" summary_model="informative", # or "conversational" ) ``` ### Sentiment Analysis `sentiment_analysis: true` **What it does:** Scores per-utterance sentiment (positive / neutral / negative). **Output:** Array of `{ text, sentiment, confidence, start, end }`. **Great for:** Customer satisfaction tracking, coaching, churn prediction. **Notes:** Segment-level (not global mood); sarcasm and very short utterances are harder to classify. **Example:** ```python for utterance in transcript.sentiment_analysis_results: if utterance.sentiment == "NEGATIVE": print(f"Negative sentiment detected: {utterance.text}") ``` ### Entity Detection `entity_detection: true` **What it does:** Extracts named entities (people, organizations, locations, products, etc.). **Output:** Array of `{ entity_type, text, start, end }`. **Great for:** Auto-tagging topics, tracking competitors mentioned, CRM enrichment. **Notes:** Operates on post-redaction text if PII redaction is enabled. **Example:** ```python # Extract all organizations mentioned organizations = [ entity.text for entity in transcript.entities if entity.entity_type == "organization" ] print(f"Companies mentioned: {', '.join(organizations)}") ``` ### Redact PII Text `redact_pii: true` **What it does:** Scans transcript for personally identifiable information and replaces matches per policy. **Output:** `text` with replacements; original `words` timing preserved. **Great for:** GDPR/CCPA compliance, safe sharing, SOC2 requirements. **Notes:** Runs **before** downstream features; they see the redacted text. **Recommended policies for meetings:** ```python config = aai.TranscriptionConfig( redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, # Remove names PIIRedactionPolicy.email_address, # Remove emails PIIRedactionPolicy.phone_number, # Remove phone numbers PIIRedactionPolicy.organization, # Remove company names ], redact_pii_sub=PIISubstitutionPolicy.hash, # Stable hash tokens ) ``` **Why hash substitution?** - Stable across the file (same value → same token) - Maintains sentence structure for LLM processing - Prevents reconstruction of original data ### Redact PII Audio `redact_pii_audio: true` **What it does:** Produces a second audio file where redacted portions are bleeped/silenced. **Output:** `redacted_audio_url` in the transcript response. **Great for:** External sharing, training materials, demos. **Notes:** Original audio is untouched; bleeped sections may sound choppy. ### Complete Example ```python expandable config = aai.TranscriptionConfig( # Core transcription speaker_labels=True, # Speech Understanding summarization=True, sentiment_analysis=True, entity_detection=True, # PII protection redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.email_address, PIIRedactionPolicy.phone_number, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True, ) transcript = transcriber.transcribe(audio_url, config=config) # Access all features meeting_insights = { "summary": transcript.summary, "sentiment_trend": analyze_sentiment_trend(transcript.sentiment_analysis_results), "entities": extract_entities(transcript.entities), "safe_transcript": transcript.text, # PII redacted "safe_audio": transcript.redacted_audio_url, # PII bleeped } ``` ## How Do I Improve the Accuracy of My Notetaker? **Best practices:** - Include participant names for better speaker recognition - Add company-specific jargon and acronyms - Include product names and technical terms - Keep individual terms under 50 characters - Up to 200 terms per request (Universal-2) or 1000 terms (Universal-3.5 Pro) ### Using Keyterms Prompt for Pre-recorded Transcription Keyterms prompting improves recognition accuracy for domain-specific vocabulary by up to 21%: ```python expandable # Define domain-specific vocabulary company_terms = [ "AssemblyAI", "Universal-3.5 Pro", "Speech Understanding", "diarization" ] participant_names = [ "Dylan Fox", "Sarah Chen", "Michael Rodriguez" ] technical_terms = [ "API endpoint", "WebSocket", "latency metrics", "TTFT" ] # Configure with keyterms prompt config = aai.TranscriptionConfig( keyterms_prompt=company_terms + participant_names + technical_terms, speaker_labels=True, # ... other settings ) ``` ### Using Keyterms Prompt for Streaming ```python expandable # Streaming with contextual keyterms keyterms = [ # Participant names "Alice Johnson", "Bob Smith", # Meeting-specific vocabulary "Q4 objectives", "revenue targets", "customer acquisition", # Technical terms "API integration", "cloud migration" ] CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "format_turns": True, "keyterms_prompt": keyterms, } ``` ## How Do I Process the Response from the API? ### Processing Pre-recorded Responses ```python expandable def process_transcript(transcript): """ Extract and process all relevant data from pre-recorded transcript """ # Basic transcript data meeting_data = { "id": transcript.id, "duration": transcript.audio_duration, "confidence": transcript.confidence, "full_text": transcript.text } # Process speaker utterances speakers = {} for utterance in transcript.utterances: speaker = utterance.speaker if speaker not in speakers: speakers[speaker] = { "utterances": [], "total_speaking_time": 0, "word_count": 0 } speakers[speaker]["utterances"].append({ "text": utterance.text, "start": utterance.start, "end": utterance.end, "confidence": utterance.confidence }) # Calculate speaking time speakers[speaker]["total_speaking_time"] += (utterance.end - utterance.start) / 1000 speakers[speaker]["word_count"] += len(utterance.text.split()) meeting_data["speakers"] = speakers # Extract summary if transcript.summary: meeting_data["summary"] = transcript.summary # Calculate meeting statistics total_duration = transcript.audio_duration # Already in seconds meeting_data["statistics"] = { "total_speakers": len(speakers), "total_words": sum(s["word_count"] for s in speakers.values()), "average_confidence": transcript.confidence, "speaking_distribution": { speaker: { "percentage": (data["total_speaking_time"] / total_duration) * 100, "minutes": data["total_speaking_time"] / 60 } for speaker, data in speakers.items() } } return meeting_data # Example usage result = process_transcript(transcript) print(f"Meeting had {result['statistics']['total_speakers']} speakers") print(f"Speaker A spoke for {result['statistics']['speaking_distribution']['A']['minutes']:.1f} minutes") ``` ### Processing Streaming Responses ```python expandable class StreamingResponseProcessor: def __init__(self): self.partial_buffer = "" self.final_transcripts = [] self.turn_metadata = [] def process_message(self, message: dict): """ Process real-time streaming messages """ msg_type = message.get("type") if msg_type == "Begin": return { "event": "session_started", "session_id": message.get("id"), "expires_at": message.get("expires_at") } elif msg_type == "Turn": return self.process_turn(message) elif msg_type == "Termination": return { "event": "session_ended", "audio_duration": message.get("audio_duration_seconds"), "session_duration": message.get("session_duration_seconds") } def process_turn(self, data: dict): """Process turn messages""" is_final = data.get("end_of_turn") transcript = data.get("transcript", "") turn_order = data.get("turn_order") response = { "turn_order": turn_order, "is_final": is_final, "confidence": data.get("end_of_turn_confidence", 0) } # Handle partials (for live display) if not is_final and transcript: self.partial_buffer = transcript response["event"] = "partial" response["text"] = transcript # Handle finals (for storage) elif is_final: final_transcript = { "turn_order": turn_order, "text": transcript, "confidence": data.get("end_of_turn_confidence"), "timestamp": datetime.now().isoformat() } self.final_transcripts.append(final_transcript) response["event"] = "final" response["text"] = transcript # Clear partial buffer self.partial_buffer = "" return response def get_full_transcript(self): """ Combine all final transcripts into complete meeting transcript """ return { "full_text": " ".join(t["text"] for t in self.final_transcripts), "transcripts": self.final_transcripts, "total_turns": len(self.final_transcripts) } # Example usage processor = StreamingResponseProcessor() # If you're using `websockets` version 13.0 or later, use `additional_headers` parameter. For older versions (< 13.0), use `extra_headers` instead. async with websockets.connect(API_ENDPOINT, additional_headers=headers) as ws: async for message in ws: data = json.loads(message) result = processor.process_message(data) if result["event"] == "partial": # Update UI with live transcript update_live_caption(result["text"]) elif result["event"] == "final": # Save final transcript save_transcript_segment(result) # Get complete transcript when done full_transcript = processor.get_full_transcript() ``` ## Additional Resources - [Universal Pre-recorded Documentation](/pre-recorded-audio) - [Universal-3.5 Pro Streaming Documentation](/streaming) - [Speaker Diarization Guide](/pre-recorded-audio/label-speakers) - [Speaker Identification Guide](/speech-understanding/speaker-identification) - [Translation Guide](/speech-understanding/translation) - [Getting Started Guide](/pre-recorded-audio/getting-started/transcribe-an-audio-file) - [API Playground](https://www.assemblyai.com/playground/streaming) - [Changelog](https://www.assemblyai.com/changelog) - [Support](https://www.assemblyai.com/contact/support) --- # Build a medical scribe URL: https://www.assemblyai.com/docs/medical-scribe-best-practices Source: docs/medical-scribe-best-practices.mdx Navigation: Overview > Use cases & integrations > Use case guides > Build a medical scribe Description: Build AI-powered medical scribes that transcribe patient encounters and generate clinical documentation. AssemblyAI provides everything you need to build a medical scribe, from high-accuracy transcription with medical terminology support to HIPAA-compliant PII redaction and structured clinical note generation through the LLM Gateway. Choose the guide that matches your clinical workflow: Transcribe recorded patient encounters with Universal-3.5 Pro. Includes Medical Mode, speaker diarization, entity detection, PII redaction, and SOAP note generation. Stream audio from a microphone during live encounters with Universal-3.5 Pro Streaming. Includes LLM Gateway post-processing, medical keyterms, and automatic SOAP note generation. ## Which approach should I use? | Scenario | Recommended approach | | --- | --- | | Post-visit documentation with highest accuracy | [Post-visit medical scribe](/medical-scribe-best-practices/medical-scribe-post-visit) | | Live encounter transcription during patient visits | [Real-time medical scribe](/medical-scribe-best-practices/medical-scribe-real-time) | | Telemedicine or emergency department visits | [Real-time medical scribe](/medical-scribe-best-practices/medical-scribe-real-time) | | Specialist consultations with complex terminology | [Post-visit medical scribe](/medical-scribe-best-practices/medical-scribe-post-visit) | | Hybrid: real-time notes with post-visit verification | Use both guides together | --- # Build a Post-Visit Medical Scribe URL: https://www.assemblyai.com/docs/medical-scribe-best-practices/medical-scribe-post-visit Source: docs/medical-scribe-best-practices/medical-scribe-post-visit.mdx Navigation: Overview > Use cases & integrations > Use case guides > Build a medical scribe Description: Complete example implementing pre-recorded transcription with Universal-3 Pro for post-visit clinical documentation. This example implements a post-visit medical scribe using pre-recorded transcription with Universal-3.5 Pro. It transcribes a clinical encounter with speaker diarization, Medical Mode, entity detection, and PII redaction, then outputs a structured provider-patient dialogue with detected clinical entities. For real-time transcription during live encounters, see the [Real-Time Medical Scribe](/medical-scribe-best-practices/medical-scribe-real-time) guide instead. ```python expandable import assemblyai as aai import asyncio from typing import Dict, List from assemblyai.types import ( PIIRedactionPolicy, PIISubstitutionPolicy, ) # Configure API key aai.settings.api_key = "your_api_key_here" async def transcribe_encounter_async(audio_source: str) -> Dict: """ Asynchronously transcribe a medical encounter with Universal-3.5 Pro Args: audio_source: Either a local file path or publicly accessible URL """ # Configure comprehensive medical transcription config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], domain="medical-v1", language_detection=True, # Diarize provider and patient speaker_labels=True, speakers_expected=2, # Typically provider and patient # Punctuation and Formatting punctuate=True, format_text=True, # Boost accuracy of medical terminology keyterms_prompt=[ # Patient-specific context "hypertension", "diabetes mellitus type 2", "metformin", # Specialty-specific terms "auscultation", "palpation", "differential diagnosis", "chief complaint", "review of systems", "physical examination", # Common medications "lisinopril", "atorvastatin", "levothyroxine", # Procedure terms "electrocardiogram", "complete blood count", "hemoglobin A1c" ], # Speech understanding for medical documentation entity_detection=True, # Extract medications, conditions, procedures redact_pii=True, # HIPAA compliance redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.date_of_birth, PIIRedactionPolicy.phone_number, PIIRedactionPolicy.email_address, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True # Create HIPAA-compliant audio ) # Create async transcriber transcriber = aai.Transcriber() try: # Submit transcription job - works with both file paths and URLs transcript = await asyncio.to_thread( transcriber.transcribe, audio_source, config=config ) # Check status if transcript.status == aai.TranscriptStatus.error: raise Exception(f"Transcription failed: {transcript.error}") # Process speaker-labeled utterances print("\n=== PROVIDER-PATIENT DIALOGUE ===\n") for utterance in transcript.utterances: # Format timestamp start_time = utterance.start / 1000 # Convert to seconds end_time = utterance.end / 1000 # Identify speaker role speaker_label = "Provider" if utterance.speaker == "A" else "Patient" # Print formatted utterance print(f"[{start_time:.1f}s - {end_time:.1f}s] {speaker_label}:") print(f" {utterance.text}") print(f" Confidence: {utterance.confidence:.2%}\n") # Extract clinical entities if transcript.entities: print("\n=== CLINICAL ENTITIES DETECTED ===\n") medications = [e for e in transcript.entities if e.entity_type == "drug"] conditions = [e for e in transcript.entities if e.entity_type == "medical_condition"] procedures = [e for e in transcript.entities if e.entity_type == "medical_process"] if medications: print("Medications:", ", ".join([m.text for m in medications])) if conditions: print("Conditions:", ", ".join([c.text for c in conditions])) if procedures: print("Procedures:", ", ".join([p.text for p in procedures])) return { "transcript": transcript, "utterances": transcript.utterances, "entities": transcript.entities, "redacted_audio_url": transcript.redacted_audio_url } except Exception as e: print(f"Error during transcription: {e}") raise async def main(): """ Example usage for medical encounter """ # Can use either local file or URL audio_source = "path/to/patient_encounter.mp3" # Or use URL # audio_source = "https://your-secure-storage.com/encounter.mp3" try: result = await transcribe_encounter_async(audio_source) # Additional processing print(f"\nEncounter duration: {result['transcript'].audio_duration} seconds") # Could send to LLM Gateway for SOAP note generation here except Exception as e: print(f"Failed to process encounter: {e}") if __name__ == "__main__": asyncio.run(main()) ``` ## Next steps - [Build a Real-Time Medical Scribe](/medical-scribe-best-practices/medical-scribe-real-time) — Stream audio from a microphone with LLM post-processing - [Medical Mode](/pre-recorded-audio/medical-mode) — Improve medical terminology accuracy - [End-to-end medical scribe pipeline](/getting-started/end-to-end-examples/medical-scribe) — Compact pipeline with SOAP note generation --- # Build a Real-Time Medical Scribe URL: https://www.assemblyai.com/docs/medical-scribe-best-practices/medical-scribe-real-time Source: docs/medical-scribe-best-practices/medical-scribe-real-time.mdx Navigation: Overview > Use cases & integrations > Use case guides > Build a medical scribe Description: Complete example for real-time streaming medical transcription with LLM Gateway post-processing and SOAP note generation. This example implements a real-time medical scribe using Universal-3.5 Pro Streaming with LLM Gateway post-processing. It uses [Medical Mode](/streaming/medical-mode) to improve accuracy for clinical terminology, streams audio from a microphone, applies LLM-powered clinical editing on each turn, and generates a SOAP note at the end of the session. For post-visit documentation using pre-recorded audio, see the [Post-Visit Medical Scribe](/medical-scribe-best-practices/medical-scribe-post-visit) guide instead. ```python expandable import os import json import time import threading from datetime import datetime from urllib.parse import urlencode import pyaudio import websocket import requests from dotenv import load_dotenv from simple_term_menu import TerminalMenu # Load environment variables from .env if present try: load_dotenv() except Exception: pass """ Medical Scribe – Real-time STT + LLM Gateway Enhancement (SOAP-ready) What this does -------------- 1) Streams mic audio to AssemblyAI Real-time STT 2) On every utterance or end of turn, calls AssemblyAI LLM Gateway to apply *medical* edits (terminology, punctuation, proper nouns, etc.) 3) Logs encounter turns and generates a SOAP note at session end via the Gateway Quick start ----------- export ASSEMBLYAI_API_KEY=your_key python medical_scribe_llm_gateway.py """ # === Config === ASSEMBLYAI_API_KEY = os.environ.get("ASSEMBLYAI_API_KEY", "your_api_key_here") # WebSocket / STT parameters - CONSERVATIVE SETTINGS FOR MEDICAL CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "domain": "medical-v1", # Enable Medical Mode for clinical terminology accuracy # MEDICAL SCRIBE CONFIGURATION - Conservative for clinical accuracy # Medical conversations have LONG pauses (provider thinking, examining patient, reviewing charts) # universal-3-5-pro defaults: min_turn_silence=100ms, max_turn_silence=1000ms "min_turn_silence": 800, # Wait much longer (vs ~100ms for voice agents, 560ms for meetings) "max_turn_silence": 2000, # Longer for clinical thinking pauses } API_ENDPOINT_BASE_URL = "wss://streaming.assemblyai.com/v3/ws" API_ENDPOINT = f"{API_ENDPOINT_BASE_URL}?{urlencode(CONNECTION_PARAMS)}" # Audio config FRAMES_PER_BUFFER = 800 # 50ms @ 16kHz SAMPLE_RATE = CONNECTION_PARAMS["sample_rate"] CHANNELS = 1 FORMAT = pyaudio.paInt16 # Globals audio = None stream = None ws_app = None audio_thread = None stop_event = threading.Event() encounter_buffer = [] # list of dicts with turn data last_processed_turn = None # === Model selection === AVAILABLE_MODELS = [ {"id": "claude-haiku-4-5-20251001", "name": "Claude Haiku 4.5", "description": "Fastest Claude, good for simple tasks"}, {"id": "claude-sonnet-4-5-20250929", "name": "Claude Sonnet 4.5", "description": "Best for coding & agents"}, {"id": "claude-sonnet-4-6", "name": "Claude Sonnet 4.6", "description": "Latest Sonnet, fast with strong reasoning"}, ] def select_model(): menu_entries = [f"{m['name']} - {m['description']}" for m in AVAILABLE_MODELS] terminal_menu = TerminalMenu( menu_entries, title="Select a model (Use ↑↓ arrows, Enter to select):", menu_cursor="❯ ", menu_cursor_style=("fg_cyan", "bold"), menu_highlight_style=("bg_cyan", "fg_black"), cycle_cursor=True, clear_screen=False, show_search_hint=True, ) idx = terminal_menu.show() if idx is None: print("Model selection cancelled. Exiting...") raise SystemExit(0) return AVAILABLE_MODELS[idx]["id"] selected_model = None # === Gateway helpers === def _gateway_chat(messages, max_tokens=800, temperature=0.2, retries=3, backoff=0.75): """Call AssemblyAI LLM Gateway with debug logging and retry.""" url = "https://llm-gateway.assemblyai.com/v1/chat/completions" headers = { "Authorization": ASSEMBLYAI_API_KEY, "Content-Type": "application/json", } payload = { "model": selected_model, "messages": messages, "max_tokens": max_tokens, "temperature": temperature, } last = None for attempt in range(retries): try: print(f"[LLM] POST {url} (model={selected_model}, attempt {attempt+1}/{retries})") resp = requests.post(url, headers=headers, json=payload, timeout=60) print(f"[LLM] ← status {resp.status_code}, bytes {len(resp.content)}") last = resp except Exception as e: if attempt == retries - 1: raise RuntimeError(f"Gateway request error: {e}") time.sleep(backoff * (attempt + 1)) continue if resp.status_code == 200: data = resp.json() if not data.get("choices") or not data["choices"][0].get("message"): raise RuntimeError(f"Gateway OK but empty body: {str(data)[:200]}") return data if resp.status_code in (429, 500, 502, 503, 504): print(f"[LLM RETRY] {resp.status_code}: {resp.text[:180]}") time.sleep(backoff * (attempt + 1)) continue raise RuntimeError(f"Gateway error {resp.status_code}: {resp.text[:300]}") raise RuntimeError( f"Gateway failed after retries. Last={getattr(last,'status_code','n/a')} {getattr(last,'text','')[:180]}" ) def post_process_with_llm(text: str) -> str: """Medical editing & normalization using LLM Gateway.""" system = { "role": "system", "content": ( "You are a clinical transcription editor. Keep the speaker's words, " "fix medical terminology (drug names, dosages, anatomy), proper nouns, " "and punctuation for readability. Preserve meaning and avoid inventing " "details. Prefer U.S. clinical style. If a medication or condition is " "phonetically close, correct to the most likely clinical term." ), } user = { "role": "user", "content": ( "Edit this short transcript for medical accuracy and readability.\n\n" f"Transcript:\n{text}" ), } try: res = _gateway_chat([system, user], max_tokens=600) return res["choices"][0]["message"]["content"].strip() except Exception as e: print(f"[LLM EDIT ERROR] {e}. Falling back to original.") return text def generate_clinical_note(): """Create a SOAP note from the encounter buffer via Gateway.""" if not encounter_buffer: print("No encounter data to summarize.") return print("\n=== GENERATING CLINICAL DOCUMENTATION (SOAP) ===") # Build a compact transcript string for the LLM lines = [] for e in encounter_buffer: if e.get("type") == "utterance": lines.append(f"[{e['timestamp']}] {e.get('speaker', 'Speaker')}: {e['text']}") elif e.get("type") == "final": lines.append(f"[{e['timestamp']}] FINAL: {e['text']}") combined = "\n".join(lines) system = { "role": "system", "content": ( "You are a clinician generating concise, structured notes. " "Produce a SOAP note (Subjective, Objective, Assessment, Plan). " "Use bullet points, keep it factual, infer reasonable clinical semantics " "from the transcript but do NOT invent data. Include medications with dosage " "and frequency if mentioned." ), } user = { "role": "user", "content": ( "Create a SOAP note from this clinical encounter transcript.\n\n" f"Transcript:\n{combined}\n\n" "Format strictly as:\n" "Subjective:\n- ...\n\nObjective:\n- ...\n\nAssessment:\n- ...\n\nPlan:\n- ...\n" ), } try: res = _gateway_chat([system, user], max_tokens=1200) soap = res["choices"][0]["message"]["content"].strip() fname = f"clinical_note_soap_{datetime.now().strftime('%Y%m%d_%H%M%S')}.txt" with open(fname, "w", encoding="utf-8") as f: f.write(soap) print(f"SOAP note saved: {fname}") except Exception as e: print(f"[SOAP ERROR] {e}") # === WebSocket callbacks === def on_open(ws): print("=" * 80) print(f"[{datetime.now().strftime('%H:%M:%S')}] Medical transcription started") print(f"Connected to: {API_ENDPOINT_BASE_URL}") print(f"Gateway model: {selected_model}") print("=" * 80) print("\nSpeak to begin. Press Ctrl+C to stop.\n") def stream_audio(): global stream while not stop_event.is_set(): try: audio_data = stream.read(FRAMES_PER_BUFFER, exception_on_overflow=False) ws.send(audio_data, websocket.ABNF.OPCODE_BINARY) except Exception as e: if not stop_event.is_set(): print(f"Error streaming audio: {e}") break global audio_thread audio_thread = threading.Thread(target=stream_audio, daemon=True) audio_thread.start() def on_message(ws, message): global last_processed_turn try: data = json.loads(message) msg_type = data.get("type") if msg_type == "Begin": print(f"[SESSION] Started - ID: {data.get('id','N/A')}\n") elif msg_type == "Turn": end_of_turn = data.get("end_of_turn", False) transcript = data.get("transcript", "") utterance = data.get("utterance", "") turn_order = data.get("turn_order", 0) # live partials if not end_of_turn and transcript: print(f"\r[PARTIAL] {transcript[:120]}...", end="", flush=True) # If AssemblyAI has finalized a turn, LLM-edit the transcript if end_of_turn and transcript: if last_processed_turn == turn_order: return # avoid duplicate processing last_processed_turn = turn_order ts = datetime.now().strftime('%H:%M:%S') print("\n[DEBUG] EOT received. Calling LLM…") edited = post_process_with_llm(transcript) changed = "(edited)" if edited.strip() != transcript.strip() else "(no change)" print(f"\n[{ts}] [FINAL {changed}]") print(f" ├─ Original STT : {transcript}") print(f" └─ Edited by LLM: {edited}") print(f"Turn: {turn_order}") encounter_buffer.append({ "timestamp": ts, "text": edited, "original_text": transcript, "turn_order": turn_order, "type": "final", }) # If we also get per-utterance chunks, just log them raw (no LLM) for timeline elif utterance: ts = datetime.now().strftime('%H:%M:%S') low = utterance.lower() if any(t in low for t in ["medication", "prescribe", "dosage", "mg", "daily"]): print(" 💊 MEDICATION MENTIONED") if any(t in low for t in ["pain", "symptom", "complaint", "problem"]): print(" 🏥 SYMPTOM REPORTED") if any(t in low for t in ["diagnose", "assessment", "impression"]): print(" 📋 DIAGNOSIS DISCUSSED") encounter_buffer.append({ "timestamp": ts, "text": utterance, "original_text": utterance, "turn_order": turn_order, "type": "utterance", }) print() elif msg_type == "Termination": dur = data.get("audio_duration_seconds", 0) print(f"\n[SESSION] Terminated – Duration: {dur}s") save_encounter_transcript() generate_clinical_note() elif msg_type == "Error": print(f"\n[ERROR] {data.get('error', 'Unknown error')}") except json.JSONDecodeError as e: print(f"Error decoding message: {e}") except Exception as e: print(f"Error handling message: {e}") def on_error(ws, error): print(f"\n[WEBSOCKET ERROR] {error}") stop_event.set() def on_close(ws, close_status_code, close_msg): print(f"\n[WEBSOCKET] Disconnected – Status: {close_status_code}") global stream, audio stop_event.set() if stream: if stream.is_active(): stream.stop_stream() stream.close() stream = None if audio: audio.terminate() audio = None if audio_thread and audio_thread.is_alive(): audio_thread.join(timeout=1.0) # === Persist artifacts === def save_encounter_transcript(): if not encounter_buffer: print("No encounter data to save.") return fname = f"encounter_transcript_{datetime.now().strftime('%Y%m%d_%H%M%S')}.txt" with open(fname, "w", encoding="utf-8") as f: f.write("Clinical Encounter Transcript\n") f.write(f"Date: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n") f.write("=" * 80 + "\n\n") for e in encounter_buffer: if e.get("speaker"): f.write(f"[{e['timestamp']}] {e['speaker']}: {e['text']}\n") else: f.write(f"[{e['timestamp']}] {e['text']}\n") f.write("\n") print(f"Encounter transcript saved: {fname}") # === Main === def run(): global audio, stream, ws_app, selected_model print("=" * 60) print(" 🎙️ Medical Scribe - STT + LLM Gateway") print("=" * 60) selected_model = select_model() print(f"✓ Using model: {selected_model}") # Init mic audio = pyaudio.PyAudio() try: stream = audio.open( input=True, frames_per_buffer=FRAMES_PER_BUFFER, channels=CHANNELS, format=FORMAT, rate=SAMPLE_RATE, ) print("Audio stream opened successfully.") except Exception as e: print(f"Error opening audio stream: {e}") if audio: audio.terminate() return # Connect WS ws_app = websocket.WebSocketApp( API_ENDPOINT, header={"Authorization": ASSEMBLYAI_API_KEY}, on_open=on_open, on_message=on_message, on_error=on_error, on_close=on_close, ) ws_thread = threading.Thread(target=ws_app.run_forever, daemon=True) ws_thread.start() try: while ws_thread.is_alive(): time.sleep(0.1) except KeyboardInterrupt: print("\n\nCtrl+C received. Stopping...") stop_event.set() # best-effort terminate if ws_app and ws_app.sock and ws_app.sock.connected: try: ws_app.send(json.dumps({"type": "Terminate"})) time.sleep(2) except Exception as e: print(f"Error sending termination: {e}") if ws_app: ws_app.close() ws_thread.join(timeout=2.0) finally: if stream and stream.is_active(): stream.stop_stream() if stream: stream.close() if audio: audio.terminate() print("Cleanup complete. Exiting.") if __name__ == "__main__": run() ``` ## Using a different streaming model If you switch to a different streaming model such as Universal-Streaming, formatting is not applied automatically. Two alternative models are available: - [`universal-streaming-english`](/streaming/getting-started/transcribe-streaming-audio) — English only - [`universal-streaming-multilingual`](/streaming/getting-started/transcribe-streaming-audio) — Supports 6 languages: English, Spanish, German, French, Italian, and Portuguese Add `format_turns=True` to your connection parameters to receive transcripts with punctuation, casing, and inverse text normalization (for example, dates, times, and phone numbers): ```python CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-streaming-english", # or "universal-streaming-multilingual" "format_turns": True, ... } ``` When `format_turns` is enabled, the model emits two `Turn` messages when a turn ends: one with `turn_is_formatted: false` (the raw unformatted transcript) and a second with `turn_is_formatted: true` (the formatted transcript). To avoid calling the LLM twice, check for both `end_of_turn` and `turn_is_formatted`: ```python elif msg_type == "Turn": end_of_turn = data.get("end_of_turn", False) turn_is_formatted = data.get("turn_is_formatted", False) transcript = data.get("transcript", "") if end_of_turn and turn_is_formatted and transcript: # Formatted final transcript — safe to post-process with LLM edited = post_process_with_llm(transcript) ... ``` Note that `turn_is_formatted` should not be used on its own to detect end of turn — always use `end_of_turn` for that. Universal-Streaming also uses a different turn detection system than U3 Pro. Instead of punctuation-based detection, it uses a confidence threshold controlled by `end_of_turn_confidence_threshold` (default `0.4`). The `min_turn_silence` and `max_turn_silence` parameters in the main code above are U3 Pro–specific and should be replaced with `end_of_turn_confidence_threshold` when switching models: ```python CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-streaming-english", "format_turns": True, "end_of_turn_confidence_threshold": 0.6, # increase for fewer false turn endings } ``` For Universal-3.5 Pro (`universal-3-5-pro`), `format_turns` is not needed since formatting is built into the end-of-turn system and `end_of_turn` and `turn_is_formatted` always have the same value. ## Next steps - [Build a Post-Visit Medical Scribe](/medical-scribe-best-practices/medical-scribe-post-visit) — Pre-recorded transcription for post-visit documentation - [Medical Mode for Streaming](/streaming/medical-mode) — Improve streaming medical terminology accuracy - [Universal-3.5 Pro Streaming](/streaming/getting-started/transcribe-streaming-audio) — Full streaming model documentation --- # Best Practices for Building Voice Agents URL: https://www.assemblyai.com/docs/voice-agents/best-practices Source: docs/voice-agents/best-practices.mdx Navigation: Overview > Use cases & integrations > Use case guides Description: Complete guide for building voice agents with AssemblyAI's Universal-3 Pro Streaming ## Introduction AssemblyAI's **Universal-3.5 Pro Streaming** is the most accurate real-time speech-to-text model designed for voice agents. It delivers formatted, immutable transcripts with sub-300ms latency, exceptional entity accuracy, native multilingual code switching, and a fully promptable interface, all optimized for conversational AI workflows. The STT component is the "ears" of your voice agent. Transcription errors propagate into the LLM and response logic, so even small accuracy gaps compound in impact. Choosing and configuring the right STT model is one of the highest-leverage decisions you can make when building a voice agent. For guidance on how to evaluate and compare STT models for your use case, see the [streaming evaluation guide](/streaming/evaluations/voice-agents). ## Why Universal-3.5 Pro Streaming for Voice Agents? Voice agents need speed, accuracy, and natural turn-taking. Universal-3.5 Pro Streaming is purpose-built for this: **Sub-300ms latency with formatted output** - Immutable transcripts arrive fully formatted (punctuation, capitalization), no waiting for a separate formatting step - Every final transcript is ready for immediate LLM processing **Exceptional entity accuracy** - Credit card numbers, phone numbers, email addresses, physical addresses, and names are transcribed with high accuracy - Short utterances like "yes", "no", "mmhmm" are handled reliably **Punctuation-based turn detection** - Turn boundaries are determined by terminal punctuation (`.` `?` `!`) combined with silence thresholds - Configurable `min_turn_silence` and `max_turn_silence` parameters let you tune responsiveness vs. accuracy - No confidence-score guessing. The model understands when a sentence is complete **Fully promptable** - Contextual `prompt` parameter for describing what the audio is about - Dynamic prompting mid-session via `UpdateConfiguration`. Adapt the model to each stage of the conversation - `keyterms_prompt` for boosting recognition of specific names, brands, and domain terms **Native multilingual support** - Supports English, Spanish, French, German, Italian, and Portuguese - Automatic code-switching between languages within a single session - Language-specific prompting for improved accuracy ## What Languages Does Universal-3.5 Pro Streaming Support? Universal-3.5 Pro Streaming supports 18 languages with automatic code-switching: - English - Spanish - German - French - Portuguese - Italian - Turkish - Dutch - Swedish - Norwegian - Danish - Finnish - Hindi - Vietnamese - Arabic - Hebrew - Japanese - Mandarin The model handles code-switching natively. Speakers can switch between supported languages mid-conversation without any configuration changes. Accuracy improves when you specify the expected language in the prompt. See [Supported languages](/streaming/getting-started/transcribe-streaming-audio) for the full language list and regional dialect reference. To bias the model toward a specific language, pass the `language_code` connection parameter — see [Language selection](/streaming/getting-started/optimizing-accuracy-and-latency#language-selection). For multilingual conversations, no configuration is needed — code-switching is native. ## How Do I Get Started? ### Complete voice agent stack AssemblyAI provides a speech-to-speech [Voice Agent API](/voice-agents/voice-agent-api) that abstracts away the complexity of a full voice agent stack: managed STT, LLM, turn detection, and TTS in a single endpoint. For a cascading architecture, AssemblyAI has the best speech-to-text model. For a complete stack, you need: 1. **Speech-to-Text (STT):** AssemblyAI Universal-3.5 Pro Streaming 2. **Large Language Model (LLM):** OpenAI, Anthropic, Google, etc. 3. **Text-to-Speech (TTS):** Rime, Cartesia, ElevenLabs, etc. 4. **Orchestration:** LiveKit, Pipecat, or custom build ### Pre-built integrations **LiveKit Agents (recommended)** LiveKit provides the fastest path to a working voice agent with AssemblyAI. See [Universal-3.5 Pro Streaming on LiveKit](/voice-agents/livekit-universal-3-5-pro) for a full guide. ```python from livekit.agents import AgentSession from livekit.plugins import assemblyai, silero session = AgentSession( stt=assemblyai.STT( model="universal-3-5-pro", min_turn_silence=100, max_turn_silence=1000, vad_threshold=0.3, ), vad=silero.VAD.load( activation_threshold=0.3, ), turn_detection="stt", min_endpointing_delay=0, ) ``` **Pipecat by Daily** Pipecat is an open-source framework for conversational AI with maximum customizability. See [Universal-3.5 Pro Streaming on Pipecat](/voice-agents/pipecat-universal-3-5-pro) for a full guide. ```python from pipecat.services.assemblyai.stt import AssemblyAISTTService from pipecat.services.assemblyai.models import AssemblyAIConnectionParams stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), connection_params=AssemblyAIConnectionParams( speech_model="universal-3-5-pro", min_turn_silence=100, max_turn_silence=1000, ), vad_force_turn_endpoint=False, # Use AssemblyAI's built-in turn detection ) ``` ### Direct WebSocket connection For custom builds, connect directly to the WebSocket API: ```python expandable import json import pyaudio import websocket import threading import time from urllib.parse import urlencode API_KEY = "YOUR_API_KEY" SAMPLE_RATE = 16000 CONNECTION_PARAMS = { "sample_rate": SAMPLE_RATE, "speech_model": "universal-3-5-pro", "min_turn_silence": 100, "max_turn_silence": 1000, } API_ENDPOINT_BASE = "wss://streaming.assemblyai.com/v3/ws" API_ENDPOINT = f"{API_ENDPOINT_BASE}?{urlencode(CONNECTION_PARAMS)}" def on_message(ws, message): data = json.loads(message) if data.get("type") == "Turn": transcript = data.get("transcript", "") end_of_turn = data.get("end_of_turn", False) if end_of_turn: # Final transcript - send to LLM print(f"Final: {transcript}") else: # Partial - can start pre-emptive LLM generation print(f"Partial: {transcript}") elif data.get("type") == "SpeechStarted": # User started speaking - handle barge-in print("Speech detected, interrupt agent if speaking") ws = websocket.WebSocketApp( API_ENDPOINT, header={"Authorization": API_KEY}, on_message=on_message, ) ``` ## How Does Turn Detection Work? Universal-3.5 Pro Streaming uses a **punctuation-based turn detection system** controlled by two parameters: | Parameter | Default | Description | | ------------------ | ---------- | ------------------------------------------------------------------------ | | `min_turn_silence` | `100` ms | Silence before a speculative end-of-turn check fires. | | `max_turn_silence` | `1000` ms | Maximum silence before forcing the turn to end. | **How it works:** 1. User speaks → audio streams to AssemblyAI 2. User pauses for `min_turn_silence` → model checks for terminal punctuation (`.` `?` `!`) 3. If terminal punctuation found → turn ends immediately with `end_of_turn: true` 4. If no terminal punctuation → partial emitted with `end_of_turn: false`, turn continues 5. If silence reaches `max_turn_silence` → turn forced to end regardless of punctuation This is different from the legacy Universal-Streaming models, which used a confidence-based `end_of_turn_confidence_threshold`. Universal-3.5 Pro Streaming does not use that parameter. Turn decisions are based on punctuation after silence thresholds. ### Configuration presets ```python # Fast - quick confirmations, IVR, yes/no questions fast_params = { "speech_model": "universal-3-5-pro", "min_turn_silence": 100, "max_turn_silence": 800, } # Balanced - most voice agent conversations (recommended) balanced_params = { "speech_model": "universal-3-5-pro", "min_turn_silence": 100, "max_turn_silence": 1000, } # Patient - entity dictation, complex instructions patient_params = { "speech_model": "universal-3-5-pro", "min_turn_silence": 200, "max_turn_silence": 2000, } ``` ### Entity splitting tradeoff Lower silence values produce faster transcripts but can split entities across turns: ```text # With (min_turn_silence=100, max_turn_silence=1000) "It's John." → turn ends (period found after 100ms pause) "Smith." → new turn "At gmail.com." → new turn # With (min_turn_silence=400, max_turn_silence=2000) "It's john.smith@gmail.com." → single turn (properly formatted) ``` For voice agents, the downstream LLM can usually piece together split entities. But if your use case involves entity extraction or alphanumeric dictation, increase `min_turn_silence` and `max_turn_silence` during those portions of the conversation using [dynamic configuration updates](#how-do-i-update-configuration-mid-session). ## How Do I Handle Barge-In and Interruptions? ### SpeechStarted events Universal-3.5 Pro Streaming emits `SpeechStarted` events when speech has been detected. `SpeechStarted` is only emitted when the model produces a transcript. This makes it a reliable signal for barge-in handling: ```json { "type": "SpeechStarted", "timestamp": 14400, "confidence": 0.79 } ``` When you receive a `SpeechStarted` event: 1. Stop TTS playback immediately 2. Switch the agent back to listening mode 3. Wait for the user's full turn to complete before responding ### Filtering backchannels Users often produce short backchannel utterances ("mhm", "yeah", "um", "okay") while the agent is speaking. Treating every `SpeechStarted` event as a barge-in causes the agent to stop mid-sentence on these fillers, even though the user didn't intend to interrupt. The fix is to gate barge-in on each `Turn` event during agent speech: suppress the interrupt when the transcript is short or every token is a known backchannel. Implementation depends on your stack: - **Voice Agent API**: semantic interruption classification is built in. See [Turn detection and interruptions](/voice-agents/voice-agent-api/turn-detection-and-interruptions). - **LiveKit**: LiveKit Cloud users should enable [adaptive interruption handling](/voice-agents/livekit-universal-3-5-pro#interruption-handling). Self-hosted deployments use the two custom filters in the same section. - **Direct WebSocket**: see [Interruption handling](/voice-agents/universal-3-5-pro-streaming-api#interruption-handling) on the Universal-3.5 Pro Streaming API page for the combined word-count + filler-word filter. ### VAD threshold alignment Universal-3.5 Pro Streaming includes an internal Silero VAD controlled by the `vad_threshold` parameter (default `0.3`). If you're also running a local VAD (common in LiveKit and Pipecat), align the thresholds to avoid a dead zone where one detects speech but the other doesn't: ```python # Both thresholds aligned at 0.3 stt = assemblyai.STT( model="universal-3-5-pro", vad_threshold=0.3, ) vad = silero.VAD.load( activation_threshold=0.3, ) ``` If you're in a noisy environment and getting false speech triggers, raise both thresholds together. ## How Can I Use Prompting to Improve Accuracy? ### The prompt parameter Universal-3.5 Pro Streaming supports a `prompt` parameter for [contextual prompting](/streaming/prompting-and-keyterms) — a natural-language description of what the audio is about. Transcription behavior itself (verbatim output, punctuation, turn detection) is built in and managed automatically; the prompt carries context, not instructions. **Beta feature** Prompting is a beta feature. We recommend starting without a prompt to establish baseline performance, then adding context to optimize for your use case. ```python CONNECTION_PARAMS = { "speech_model": "universal-3-5-pro", "prompt": "AI voice agent call with a customer about an internet service outage." } ``` **Tips for effective prompts:** - **Describe the conversation**: domain, scenario, or full details — start broad, and add only details your application actually knows (see the [three context levels](/streaming/prompting-and-keyterms#contextual-prompting)) - **Include names and identifiers you already know**: caller name, account or order IDs, products — detailed context helps the model spell them correctly - **Specify language**: use the `language_code` connection parameter for monolingual sessions ([Language selection](/streaming/getting-started/optimizing-accuracy-and-latency#language-selection)) ### Keyterms prompting Use `keyterms_prompt` to boost recognition of specific names, brands, or domain terms, up to 100 terms per session: ```python CONNECTION_PARAMS = { "speech_model": "universal-3-5-pro", "keyterms_prompt": json.dumps([ "AssemblyAI", "LiveKit", "Dr. Rodriguez", "Lisinopril", "iPhone 15 Pro", ]) } ``` **Best practices for keyterms:** - Include proper names, product names, technical terms, and domain-specific jargon - Include terms up to 50 characters each - Don't include common English words, single letters, or generic phrases - Don't exceed 100 terms total For detailed guidance, see [Keyterms prompting](/streaming/prompting-and-keyterms). ## How Do I Update Configuration Mid-Session? You can update `prompt`, `keyterms_prompt`, `min_turn_silence`, and `max_turn_silence` during an active session using `UpdateConfiguration`. This is one of Universal-3.5 Pro Streaming's most powerful features for voice agents. ### Dynamic keyterms by conversation stage As your voice agent moves through different stages, update keyterms to match what the user is likely to say: ```python # Caller identification stage ws.send(json.dumps({ "type": "UpdateConfiguration", "keyterms_prompt": ["Kelly Byrne-Donoghue", "date of birth", "January", "February"] })) # Medical intake stage ws.send(json.dumps({ "type": "UpdateConfiguration", "keyterms_prompt": ["cardiology", "echocardiogram", "Dr. Patel", "metoprolol"] })) # Payment stage - also increase max_turn_silence for credit card dictation ws.send(json.dumps({ "type": "UpdateConfiguration", "keyterms_prompt": ["Visa", "Mastercard", "American Express"], "max_turn_silence": 3000 })) ``` ### Dynamic prompting You can also update the transcription prompt mid-session. This is especially powerful when paired with tool calls in your LLM: - If your agent asks a yes/no question, prompt the model to anticipate short responses - If your agent asks for a phone number or email, prompt it to expect those formats - If you present a list of options, boost those options in the prompt ```python # After asking "Would you like to confirm your appointment?" ws.send(json.dumps({ "type": "UpdateConfiguration", "prompt": "User is responding yes or no to a confirmation question. Expect short responses." })) # After asking "What's your phone number?" ws.send(json.dumps({ "type": "UpdateConfiguration", "prompt": "User is dictating a phone number. Expect digits and formatting.", "max_turn_silence": 3000 })) ``` ## How Do I Use Speaker Diarization? Streaming Diarization identifies and labels individual speakers in real time. Each `Turn` event includes a `speaker_label` field (e.g., `"A"`, `"B"`) indicating which speaker produced that transcript. Enable it by adding `speaker_labels: true` to your connection parameters: ```python CONNECTION_PARAMS = { "speech_model": "universal-3-5-pro", "speaker_labels": True, } ``` Speaker accuracy improves over the course of a session as the model accumulates embedding context. **With LiveKit:** ```python stt = assemblyai.STT( model="universal-3-5-pro", speaker_labels=True, ) ``` **With Pipecat** (including custom formatting): ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), connection_params=AssemblyAIConnectionParams( speech_model="universal-3-5-pro", speaker_labels=True, ), speaker_format="[{speaker}] {text}", ) ``` For more details, see [Streaming Diarization and Multichannel](/streaming/label-speakers-and-separate-channels). ## How Do I Optimize for Latency? ### Key optimizations **1. Use the right silence thresholds** Start with `min_turn_silence=100` and `max_turn_silence=1000`. Only increase if you're seeing entity splitting issues. **2. Tune `interruption_delay` for faster TTFT** The `interruption_delay` parameter controls how soon the first partial is emitted. Set `interruption_delay=0` for the fastest possible time to first token (~300ms effective). The default of 500ms produces a first partial at ~800ms. See [Tuning early partial timing](/streaming/getting-started/transcribe-streaming-audio) for details. **3. Eliminate additive delays in your orchestrator** In LiveKit with `turn_detection="stt"`, set `min_endpointing_delay=0`. LiveKit's default 0.5s delay is additive on top of AssemblyAI's own endpointing. **4. Use 16kHz sample rate** This balances audio quality and bandwidth. Higher sample rates don't improve accuracy. **5. Align VAD thresholds** Mismatched VAD thresholds between your local VAD and AssemblyAI create a dead zone that delays interruption. Set both to `0.3`. **6. Skip unnecessary features** Only enable `speaker_labels` if you need diarization. Only use `keyterms_prompt` if you have domain-specific terms. Each feature adds marginal processing overhead. ### Latency breakdown | Stage | Typical latency | Notes | | ----- | --------------- | ----- | | Audio to AssemblyAI | ~50ms | Network dependent | | Speech-to-text | ~200-300ms | Sub-300ms P50 | | `min_turn_silence` check | 100ms+ | Configurable | | `max_turn_silence` fallback | 1000ms+ | Only if no terminal punctuation | ## How Does the Message Sequence Work? Universal-3.5 Pro Streaming sends messages in a specific sequence. Here's what a typical conversation looks like: **1. Session begins** ```json { "type": "Begin", "id": "session-id", "expires_at": 1759796682 } ``` **2. Speech detected** ```json { "type": "SpeechStarted", "timestamp": 1200, "confidence": 0.85 } ``` **3. Early partial** (emitted after 750ms of continuous speech) ```json { "type": "Turn", "turn_order": 0, "end_of_turn": false, "turn_is_formatted": false, "transcript": "Yeah my credit card..." } ``` **4. Silence-based partial** (speaker pauses, no terminal punctuation) ```json { "type": "Turn", "turn_order": 0, "end_of_turn": false, "turn_is_formatted": false, "transcript": "Yeah my credit card number is--" } ``` **5. Final transcript** (terminal punctuation found, or `max_turn_silence` reached) ```json { "type": "Turn", "turn_order": 0, "end_of_turn": true, "turn_is_formatted": true, "transcript": "Yeah, my credit card number is 8888-8888-8888-8888.", "speaker_label": "A" } ``` For Universal-3.5 Pro Streaming, `end_of_turn` and `turn_is_formatted` always have the same value. You can reliably use `end_of_turn: true` to detect a formatted, final transcript. **6. Session termination** ```json { "type": "Termination", "audio_duration_seconds": 45.2 } ``` For the complete message reference, see [Message sequence](/streaming/getting-started/transcribe-streaming-audio). ## How Can I Improve Accuracy? ### Keyterms prompting The single most effective way to improve accuracy on domain-specific terms. Keyterms are especially useful for improving recognition of proper nouns, product names, and technical jargon spoken with accents or in noisy environments. See [How Can I Use Prompting to Improve Accuracy?](#how-can-i-use-prompting-to-improve-accuracy) above. ### Dynamic configuration updates Update keyterms and prompts mid-session based on conversation context. See [How Do I Update Configuration Mid-Session?](#how-do-i-update-configuration-mid-session) above. ### Tune silence thresholds If entities are splitting across turns, increase `min_turn_silence` (for punctuation-triggered splits) or `max_turn_silence` (for forced timeout splits). You can do this dynamically mid-session for specific conversation stages like entity dictation. ### Noise handling Universal-3.5 Pro Streaming handles background noise well out of the box. Avoid adding noise cancellation as a preprocessing step. The artifacts it introduces typically cause more harm than the background noise itself. For telephony environments with low-quality audio (such as 8 kHz mulaw), you can prompt the model to tag genuinely unclear segments as `[unclear]` rather than forcing a guess. This helps you identify audio segments that no model (or human) can reliably transcribe, and prevents inaccurate guesses from entering your downstream pipeline. ## Scaling and Rate Limits Universal-3.5 Pro Streaming provides unlimited parallel streams: - No hard caps on simultaneous connections - No overage fees for spike traffic - Automatic scaling from 5 to 50,000+ streams **Rate limits:** - Free users: 5 new streams per minute - Pay-as-you-go: 100 new streams per minute - When using 70%+ of your limit, capacity automatically increases 10% every 60 seconds These limits are designed to never interfere with legitimate applications. Your baseline limit is guaranteed and never decreases, so you can scale smoothly without artificial barriers. --- ## Evaluating your voice agent Benchmark scores are a useful starting point, but they don't tell the full story. To determine which STT model works best for your voice agent in production: - **Run A/B tests at scale.** Swap models in your agent pipeline, compare core outcome metrics (task completion rate, booking rate, resolution rate), and let real user behavior determine the winner. - **Optimize for outcomes, not benchmarks.** The question is not which model has the best WER. It is which model drives users to complete their goal most reliably. - **Simulate real scenarios early.** Before you have real users, manually test the full agent flow under realistic conditions. This surfaces issues that isolated STT benchmarks will miss. For a complete evaluation framework including accuracy metrics, latency metrics, and ground truth best practices, see the [streaming evaluation guide](/streaming/evaluations/voice-agents). ## Additional Resources - [Streaming evaluation guide](/streaming/evaluations/voice-agents) - [Universal-3.5 Pro Streaming overview](/streaming/getting-started/transcribe-streaming-audio) - [Turn detection guide](/streaming/getting-started/transcribe-streaming-audio) - [Prompting guide](/streaming/prompting-and-keyterms) - [Keyterms prompting](/streaming/prompting-and-keyterms) - [Message sequence](/streaming/getting-started/transcribe-streaming-audio) - [Streaming diarization and multichannel](/streaming/label-speakers-and-separate-channels) - [Supported languages](/streaming/getting-started/transcribe-streaming-audio) - [LiveKit integration](/voice-agents/livekit-universal-3-5-pro) - [Pipecat integration](/voice-agents/pipecat-universal-3-5-pro) - [API Playground](https://www.assemblyai.com/playground/streaming) - [Changelog](https://www.assemblyai.com/changelog) --- # Best Practices for building Contact Center Applications URL: https://www.assemblyai.com/docs/contact-center-best-practices Source: docs/contact-center-best-practices.mdx Navigation: Overview > Use cases & integrations > Use case guides Description: Complete guide for building contact center applications with AssemblyAI ## Introduction Building a contact center application requires careful consideration of accuracy, speaker separation, compliance, and scalability. This guide addresses common questions and provides practical solutions for both post-call analytics and real-time agent assist scenarios. ## Why AssemblyAI for contact centers? AssemblyAI stands out as the premier choice for contact center applications with several key advantages: ### Industry-leading accuracy on telephony audio - **Universal-3.5 Pro model** delivers best-in-class accuracy on 8kHz telephony audio - **2.9% speaker diarization error rate** for precise agent vs. customer attribution - **Multichannel support** for stereo call recordings where agent and customer are on separate channels - **Keyterms prompt** allows providing call context to improve accuracy of company names, products, and compliance phrases ### Streaming with Universal-3.5 Pro For real-time agent assist, AssemblyAI's Universal-3.5 Pro Streaming model (`universal-3-5-pro`) offers: - **Low latency** enables live transcription during calls - **Format turns** feature provides structured, readable output - **Dynamic prompting** via `UpdateConfiguration` to update context mid-call - **Dual-channel streaming** for separate agent and customer audio streams ### End-to-end voice AI platform Unlike fragmented solutions, AssemblyAI provides a unified API for: - Transcription with speaker diarization (agent vs. customer) - Multichannel audio support for stereo call recordings - PII redaction on both text and audio for HIPAA and PCI compliance - Post-processing workflows with custom prompting - from call summaries to QA scoring - Streaming and pre-recorded transcription in a single platform - Compliance and security built for enterprise workloads (BAA, SOC2, ISO) ## When should I use pre-recorded vs streaming for contact centers? Understanding when to use pre-recorded versus streaming is critical for contact center workflows. ### Pre-recorded Speech-to-text **Post-call analytics** - Call already happened, you have the full recording - **Highest accuracy needed** - Pre-recorded models have the highest accuracy - **Speaker diarization is critical** - Pre-recorded has 2.9% speaker error rate - **Multichannel recordings** - Most contact center recordings are stereo with agent and customer on separate channels - **Compliance workflows** - Full PII redaction with audio de-identification - **Post-call analytics** - Summarization, sentiment analysis, entity detection, QA scoring - **Batch processing** - Processing large volumes of call recordings **Best for:** QA scoring, compliance monitoring, coaching insights, post-call CRM updates, searchable call archives ### Streaming Speech-to-text **Live calls** - Transcribing as the call happens You should use streaming when you need to display a live transcript to agents during calls. With Universal-3.5 Pro Streaming, accuracy is closer to pre-recorded, but pre-recorded will always be the most accurate option. - **Agent assist** - Live transcription visible to agents during calls - **Real-time coaching** - Prompt agents with suggested responses or compliance reminders - **Live compliance monitoring** - Detect compliance violations in real-time - **No recording available** - Processing live audio only **Best for:** Agent assist, real-time coaching, live compliance monitoring, live call transcription ### Hybrid approach (recommended) Many contact center platforms use **both**: 1. **Streaming during the call** - Provide live transcription for agent assist and real-time coaching 2. **Pre-recorded after the call** - Generate high-quality transcript with speaker labels, summary, and analytics **Example workflow:** - Call begins → Start streaming for live agent assist - Call ends → Upload recording to pre-recorded API for final transcript with speaker names - Generate call summary, QA score, and compliance report from pre-recorded transcript - Push results to CRM (e.g., Salesforce) ## What languages and features for a contact center application? ### Pre-recorded calls (Universal-3.5 Pro) For post-call analytics, AssemblyAI supports: **Languages**: - 99 languages supported - Automatic Language Detection to route to the most spoken language - Code Switching to preserve changes in speech between languages **Core Features**: - Speaker diarization (agent-customer separation) - Multichannel audio support - when agent and customer are on separate audio channels, enables perfect speaker separation without diarization - Automatic formatting, punctuation, and capitalization - Keyterms prompting for boosting domain-specific terms (up to 1000 terms for Universal-3.5 Pro) - Natural language prompting (Universal-3.5 Pro) - up to 1,500 words to guide transcription behavior - Speaker options with configurable min/max expected speakers for call transfers **Speech Understanding**: - Summarization for call recaps - Sentiment analysis for customer satisfaction tracking - Entity detection for extracting names, account numbers, and products - Speaker identification to map generic labels to agent and customer names - Translation between 86 languages **Guardrails**: - PII redaction on text and audio for HIPAA and PCI compliance ### Streaming (Universal-3.5 Pro Streaming) For live call transcription, use **Universal-3.5 Pro Streaming** (`universal-3-5-pro`) for the highest streaming accuracy: **Core Features**: - Speaker diarization for identifying agent vs. customer - Partial and final transcripts for responsive UI - Format turns for structured, readable output - Keyterms prompt for company names, products, and compliance phrases - Dual-channel streaming for separate agent and customer audio For more details, see the [Universal-3.5 Pro Streaming documentation](/streaming/getting-started/transcribe-streaming-audio). ## How can I get started building a post-call analytics pipeline? Here's a complete example implementing pre-recorded transcription for contact center call analysis: ```python expandable import assemblyai as aai import asyncio from typing import Dict, List from assemblyai.types import ( SpeakerOptions, PIIRedactionPolicy, PIISubstitutionPolicy, ) # Configure API key aai.settings.api_key = "your_api_key_here" async def transcribe_call(audio_source: str, agent_name: str = None) -> Dict: """ Transcribe a contact center call recording with full analytics Args: audio_source: Either a local file path or publicly accessible URL agent_name: Optional agent name for speaker identification """ # Configure comprehensive call analysis config = aai.TranscriptionConfig( # Model selection speech_models=["universal-3-5-pro", "universal-2"], # Speaker diarization speaker_labels=True, speaker_options=SpeakerOptions( min_speakers_expected=2, # Agent and customer max_speakers_expected=5 # Allow for call transfers - safe to keep high ), multichannel=False, # Set to True if audio has separate channel per speaker # Language detection language_detection=True, # Boost accuracy of contact center vocabulary keyterms_prompt=[ # Company-specific terms "Acme Corp", "Premium Support Plan", # Compliance phrases "recorded line", "calls are monitored and recorded", # Common contact center terms "account number", "case number", "ticket number", "escalation", "supervisor", "hold time", ], # Post-call analytics summarization=True, sentiment_analysis=True, entity_detection=True, # PII protection for compliance redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.phone_number, PIIRedactionPolicy.email_address, PIIRedactionPolicy.account_number, PIIRedactionPolicy.us_social_security_number, PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.credit_card_cvv, PIIRedactionPolicy.credit_card_expiration, PIIRedactionPolicy.date_of_birth, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True, ) # Add speaker identification if agent name is known if agent_name: config.speech_understanding = { "request": { "speaker_identification": { "speaker_type": "role", "speakers": [ {"role": "Agent", "name": agent_name}, {"role": "Customer"} ] } } } # Create transcriber transcriber = aai.Transcriber() try: # Submit transcription job transcript = await asyncio.to_thread( transcriber.transcribe, audio_source, config=config ) # Check status if transcript.status == aai.TranscriptStatus.error: raise Exception(f"Transcription failed: {transcript.error}") # Process speaker-labeled utterances for utterance in transcript.utterances: start_time = utterance.start / 1000 # Convert ms to seconds end_time = utterance.end / 1000 print(f"[{start_time:.1f}s - {end_time:.1f}s] {utterance.speaker}:") print(f" {utterance.text}\n") return { "transcript": transcript, "utterances": transcript.utterances, "summary": transcript.summary, "sentiment": transcript.sentiment_analysis_results, "entities": transcript.entities, "redacted_audio_url": transcript.redacted_audio_url, } except Exception as e: print(f"Error during transcription: {e}") raise async def main(): audio_source = "https://your-storage.com/calls/call_recording.mp3" result = await transcribe_call(audio_source, agent_name="Sarah Johnson") print(f"\nCall duration: {result['transcript'].audio_duration} seconds") print(f"Summary: {result['summary']}") if __name__ == "__main__": asyncio.run(main()) ``` ## How Do I Handle Multichannel Contact Center Audio? Most contact center recordings are stereo with the agent on one channel and the customer on the other. Multichannel transcription gives you perfect speaker separation without diarization. ### Pre-recorded Multichannel ```python expandable config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], multichannel=True, # Enable when agent and customer are on separate channels speaker_labels=False, # Disable - channels already separate speakers # Still enable analytics summarization=True, sentiment_analysis=True, entity_detection=True, # PII redaction redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.us_social_security_number, PIIRedactionPolicy.account_number, ], redact_pii_sub=PIISubstitutionPolicy.hash, ) transcriber = aai.Transcriber() transcript = transcriber.transcribe(audio_file, config=config) # Channel 1 = Agent, Channel 2 = Customer (typical layout) for utterance in transcript.utterances: role = "Agent" if utterance.channel == "1" else "Customer" print(f"{role}: {utterance.text}") ``` **When to use multichannel:** - Call recordings from PBX systems with separate agent/customer channels - Recordings from platforms like Genesys, Twilio, Five9, NICE, or Talkdesk - Any stereo recording where each channel represents a different speaker **Benefits:** - **Perfect speaker separation** - No diarization errors - **No speaker confusion or overlap issues** - **Higher accuracy** - Model processes clean single-speaker audio per channel ### Streaming Multichannel For real-time dual-channel transcription, create separate streaming sessions per channel: ```python expandable import asyncio import websockets import json from urllib.parse import urlencode API_KEY = "your_api_key" class ChannelTranscriber: def __init__(self, channel_id: int, role: str): self.channel_id = channel_id self.role = role self.connection_params = { "sample_rate": 8000, # Telephony standard "speech_model": "universal-3-5-pro", "format_turns": True, "encoding": "pcm_mulaw", # Common telephony encoding } async def transcribe_channel(self, audio_stream): url = f"wss://streaming.assemblyai.com/v3/ws?{urlencode(self.connection_params, doseq=True)}" # If using websockets >= 13.0, use additional_headers. For < 13.0, use extra_headers. async with websockets.connect(url, additional_headers={"Authorization": API_KEY}) as ws: # Send and receive must run concurrently for real-time streaming async def send_audio(): async for audio_chunk in audio_stream: await ws.send(audio_chunk) async def receive_transcripts(): async for message in ws: data = json.loads(message) if data.get("type") == "Turn" and data.get("end_of_turn"): print(f"{self.role}: {data['transcript']}") await asyncio.gather(send_audio(), receive_transcripts()) # Create transcriber for each channel async def transcribe_live_call(agent_audio_stream, customer_audio_stream): agent = ChannelTranscriber(0, "Agent") customer = ChannelTranscriber(1, "Customer") await asyncio.gather( agent.transcribe_channel(agent_audio_stream), customer.transcribe_channel(customer_audio_stream), ) ``` See our [multichannel streaming guide](/streaming/label-speakers-and-separate-channels#multichannel-streaming-audio) for complete implementation details. ## How Can I Build a Real-Time Agent Assist? Here's a complete example for real-time streaming transcription optimized for contact center agent assist: ```python expandable # pip install pyaudio websocket-client import pyaudio import websocket import json import threading import time from urllib.parse import urlencode from datetime import datetime # --- Configuration --- YOUR_API_KEY = "your_api_key" # Contact center keyterms KEYTERMS = [ # Company and product terms "Acme Corp", "Premium Support Plan", "Enterprise License", # Compliance phrases "recorded line", "calls are monitored", # Common contact center vocabulary "account number", "case number", "escalation", "supervisor", ] # CONTACT CENTER CONFIGURATION CONNECTION_PARAMS = { "sample_rate": 8000, # Telephony standard (8kHz) "speech_model": "universal-3-5-pro", # Universal-3.5 Pro Streaming for highest accuracy "format_turns": True, # Contact center turn detection # universal-3-5-pro defaults: min_turn_silence=100ms, max_turn_silence=1000ms "min_turn_silence": 400, # Longer than default for natural call pauses "max_turn_silence": 1500, # Longer for customers explaining issues # Keyterms for accuracy "keyterms_prompt": KEYTERMS, } API_ENDPOINT_BASE_URL = "wss://streaming.assemblyai.com/v3/ws" API_ENDPOINT = f"{API_ENDPOINT_BASE_URL}?{urlencode(CONNECTION_PARAMS, doseq=True)}" # Audio Configuration FRAMES_PER_BUFFER = 400 # 50ms of audio at 8kHz SAMPLE_RATE = CONNECTION_PARAMS["sample_rate"] CHANNELS = 1 FORMAT = pyaudio.paInt16 # Global variables audio = None stream = None ws_app = None audio_thread = None stop_event = threading.Event() transcript_buffer = [] def on_open(ws): print("=" * 80) print(f"[{datetime.now().strftime('%H:%M:%S')}] Agent assist transcription started") print(f"Connected to: {API_ENDPOINT_BASE_URL}") print(f"Keyterms configured: {', '.join(KEYTERMS[:5])}...") print("=" * 80) def stream_audio(): global stream while not stop_event.is_set(): try: audio_data = stream.read(FRAMES_PER_BUFFER, exception_on_overflow=False) ws.send(audio_data, websocket.ABNF.OPCODE_BINARY) except Exception as e: if not stop_event.is_set(): print(f"Error streaming audio: {e}") break global audio_thread audio_thread = threading.Thread(target=stream_audio) audio_thread.daemon = True audio_thread.start() def on_message(ws, message): try: data = json.loads(message) msg_type = data.get("type") if msg_type == "Begin": session_id = data.get("id", "N/A") print(f"[SESSION] Started - ID: {session_id}\n") elif msg_type == "Turn": end_of_turn = data.get("end_of_turn", False) transcript = data.get("transcript", "") turn_order = data.get("turn_order", 0) # Show partials for responsive agent UI if not end_of_turn and transcript: print(f"\r[LIVE] {transcript}", end="", flush=True) # Use formatted finals for agent display if end_of_turn and transcript: timestamp = datetime.now().strftime('%H:%M:%S') print(f"\n[{timestamp}] {transcript}") # Detect compliance keywords transcript_lower = transcript.lower() if any(term in transcript_lower for term in ["cancel", "refund", "complaint", "supervisor"]): print(" ** ESCALATION KEYWORD DETECTED **") transcript_buffer.append({ "timestamp": timestamp, "text": transcript, "turn_order": turn_order, "type": "final" }) print() elif msg_type == "Termination": audio_duration = data.get("audio_duration_seconds", 0) print(f"\n[SESSION] Terminated - Duration: {audio_duration}s") elif msg_type == "Error": error_msg = data.get("error", "Unknown error") print(f"\n[ERROR] {error_msg}") except json.JSONDecodeError as e: print(f"Error decoding message: {e}") except Exception as e: print(f"Error handling message: {e}") def on_error(ws, error): print(f"\n[WEBSOCKET ERROR] {error}") stop_event.set() def on_close(ws, close_status_code, close_msg): print(f"\n[WEBSOCKET] Disconnected - Status: {close_status_code}, Message: {close_msg}") global stream, audio stop_event.set() if stream: if stream.is_active(): stream.stop_stream() stream.close() stream = None if audio: audio.terminate() audio = None if audio_thread and audio_thread.is_alive(): audio_thread.join(timeout=1.0) def run(): global audio, stream, ws_app audio = pyaudio.PyAudio() try: stream = audio.open( input=True, frames_per_buffer=FRAMES_PER_BUFFER, channels=CHANNELS, format=FORMAT, rate=SAMPLE_RATE, ) except Exception as e: print(f"Error opening audio stream: {e}") if audio: audio.terminate() return ws_app = websocket.WebSocketApp( API_ENDPOINT, header={"Authorization": YOUR_API_KEY}, on_open=on_open, on_message=on_message, on_error=on_error, on_close=on_close, ) ws_thread = threading.Thread(target=ws_app.run_forever) ws_thread.daemon = True ws_thread.start() try: while ws_thread.is_alive(): time.sleep(0.1) except KeyboardInterrupt: print("\n\nCtrl+C received. Stopping transcription...") stop_event.set() if ws_app and ws_app.sock and ws_app.sock.connected: try: terminate_message = {"type": "Terminate"} ws_app.send(json.dumps(terminate_message)) time.sleep(1) except Exception as e: print(f"Error sending termination message: {e}") if ws_app: ws_app.close() ws_thread.join(timeout=2.0) finally: if stream and stream.is_active(): stream.stop_stream() if stream: stream.close() if audio: audio.terminate() print("Cleanup complete. Exiting.") if __name__ == "__main__": run() ``` ## How Should I Handle Pre-recorded Transcription in Production? ### Webhook Callbacks (Recommended) For high-volume contact center workloads, use webhooks instead of polling: ```python expandable config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], webhook_url="https://your-app.com/webhooks/assemblyai", webhook_auth_header_name="X-Webhook-Secret", webhook_auth_header_value="your_secret_here", speaker_labels=True, multichannel=True, summarization=True, sentiment_analysis=True, entity_detection=True, redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.us_social_security_number, PIIRedactionPolicy.account_number, ], redact_pii_sub=PIISubstitutionPolicy.hash, ) # Submit job and return immediately (non-blocking) transcript = transcriber.submit(audio_url, config=config) print(f"Job submitted: {transcript.id}") # Your app continues processing other calls ``` **Webhook handler example:** ```python expandable from flask import Flask, request, jsonify app = Flask(__name__) @app.route("/webhooks/assemblyai", methods=["POST"]) def assemblyai_webhook(): if request.headers.get("X-Webhook-Secret") != "your_secret_here": return jsonify({"error": "Unauthorized"}), 401 import requests as http_requests data = request.json transcript_id = data["transcript_id"] status = data["status"] if status == "completed": # Fetch the full transcript (webhook only sends transcript_id and status) transcript = http_requests.get( f"https://api.assemblyai.com/v2/transcript/{transcript_id}", headers={"authorization": "your_api_key"} ).json() process_completed_call(transcript) elif status == "error": log_transcription_error(transcript_id) return jsonify({"received": True}), 200 def process_completed_call(transcript): """Process completed call transcript and push to CRM""" utterances = transcript["utterances"] summary = transcript["summary"] # Store in database save_to_database(transcript) # Push summary to CRM push_to_crm(transcript["id"], summary) # Run QA scoring qa_score = score_call_quality(utterances) save_qa_score(transcript["id"], qa_score) ``` ### Scaling Considerations - **HTTP rate limit:** 20,000 requests per 5-minute window, counted across submissions (POST) and polling (GET) combined - **Exceeding the limit:** returns a `403` response - **Parallel transcriptions (rate limit):** 200+ for paid accounts (queued beyond that) - **Ramp up gradually:** start at 10-50 parallel requests, double incrementally - **Avoid the rate limit:** use [webhooks](/pre-recorded-audio/webhooks) or jittered, widened polling — see [Polling without exceeding the rate limit](/pre-recorded-audio/guides/bulk-transcription-and-load-tests-at-scale#polling-without-exceeding-the-rate-limit) - **Contact Sales** before large-scale rollouts ## How Do I Handle PII and Compliance? PII redaction is critical for contact center compliance (HIPAA, PCI-DSS, GDPR, CCPA). ### Recommended PII Configuration ```python expandable config = aai.TranscriptionConfig( redact_pii=True, redact_pii_policies=[ # Customer identity PIIRedactionPolicy.person_name, PIIRedactionPolicy.date_of_birth, PIIRedactionPolicy.us_social_security_number, # Contact information PIIRedactionPolicy.phone_number, PIIRedactionPolicy.email_address, PIIRedactionPolicy.location, # Financial information (PCI-DSS) PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.credit_card_cvv, PIIRedactionPolicy.credit_card_expiration, PIIRedactionPolicy.account_number, PIIRedactionPolicy.banking_information, ], redact_pii_sub=PIISubstitutionPolicy.hash, # Stable hash tokens redact_pii_audio=True, # Create de-identified audio file ) ``` **Why hash substitution?** - Stable across the file (same value = same token) - Maintains sentence structure for downstream LLM processing - Prevents reconstruction of original data ### HIPAA Compliance - AssemblyAI provides a **Business Associate Agreement (BAA)** at no cost - Paid customers can review and sign our standard online BAA self-serve from the [**Data Controls** page](https://www.assemblyai.com/dashboard/settings/data-controls) in the dashboard (Owner or Admin role required, upgraded account required) - Use PII redaction with audio de-identification for full compliance ## How Do I Improve the Accuracy of My Contact Center Transcription? ### Prompting Best Practices The most impactful lever for contact center accuracy is **prompting**. Use a structured prompt with a `Context:` field: ```python expandable config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], # Natural language prompt for transcription guidance prompt=( "Transcribe this audio with perfect punctuation and formatting. " "Preserve linguistic speech patterns including disfluencies, filler words, " "hesitations, repetitions, stutters, false starts, and colloquialisms. " "Transcribe in the original language mix (code-switching), preserving the " "words in the language they are spoken. Output plain transcript text only. " "Use a new line when the voice changes; each line contains only one " "person's words.\n\n" "Context: Acme Corp customer service call, recorded line, " "Agent: Sarah Johnson, calls are monitored and recorded" ), # Keyterms for proper nouns and domain vocabulary keyterms_prompt=[ "Acme Corp", "Sarah Johnson", "Premium Support Plan", "Enterprise License", "recorded line", "calls are monitored and recorded", ], speaker_labels=True, ) ``` **Tips for effective prompting:** - Use **positive instructions** ("transcribe verbatim") not negative ("do NOT summarize") - **Start with fewer instructions, add one at a time** — every added instruction risks conflicting with another. Treat the older "3–6 instructions" guidance as an upper bound, not a target. - **Layer instructions one by one** and test each against your call recordings to measure impact - Dynamize the `Context:` line per call with known info: company name, agent name, compliance phrases - Use **keyterms** for proper nouns and domain vocabulary (company names, product names, agent names) ### Using Keyterms for Pre-recorded Transcription ```python expandable # Build keyterms dynamically per call call_keyterms = [ # Company terms (static) "Acme Corp", "Premium Support Plan", # Agent name (from routing system) agent_name, # Customer name (from CRM lookup) customer_name, # Account-specific terms "account ending in 4532", ] config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], keyterms_prompt=call_keyterms, speaker_labels=True, ) ``` ### Using Keyterms for Streaming ```python # Streaming with contact center context keyterms = [ "Acme Corp", "Premium Support Plan", "Sarah Johnson", "recorded line", ] CONNECTION_PARAMS = { "sample_rate": 8000, "speech_model": "universal-3-5-pro", "format_turns": True, "encoding": "pcm_mulaw", "keyterms_prompt": keyterms, } ``` ## What Workflows Can I Build for My Contact Center Application? Use these features to transform raw call transcripts into actionable insights. ### Summarization `summarization: true` **What it does:** Generates an abstractive recap of the call. **Output:** `summary` string (bullets/paragraph format). **Great for:** Post-call CRM updates, call recaps, supervisor review. ```python config = aai.TranscriptionConfig( summarization=True, summary_type="bullets", # or "bullets_verbose", "gist", "headline", "paragraph" summary_model="informative", # or "conversational" ) ``` ### Sentiment Analysis `sentiment_analysis: true` **What it does:** Scores per-utterance sentiment (positive / neutral / negative). **Output:** Array of `{ text, sentiment, confidence, start, end }`. **Great for:** Customer satisfaction tracking, escalation detection, QA scoring. ```python # Analyze customer sentiment across a call negative_count = 0 for result in transcript.sentiment_analysis_results: if result.sentiment == "NEGATIVE": negative_count += 1 print(f"Negative at {result.start / 1000:.1f}s: {result.text}") # Flag calls with high negative sentiment if negative_count > 3: flag_for_supervisor_review(transcript.id) ``` ### Entity Detection `entity_detection: true` **What it does:** Extracts named entities (people, organizations, locations, products, etc.). **Output:** Array of `{ entity_type, text, start, end }`. **Great for:** CRM enrichment, auto-tagging topics, competitor tracking. ```python # Extract key entities from a call organizations = [e.text for e in transcript.entities if e.entity_type == "organization"] print(f"Companies mentioned: {', '.join(organizations)}") ``` ### Speaker Identification Map generic speaker labels to agent and customer names: ```python config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], speaker_labels=True, speech_understanding={ "request": { "speaker_identification": { "speaker_type": "role", "speakers": [ {"role": "Agent", "name": "Sarah Johnson", "description": "Customer service representative"}, {"role": "Customer"} ] } } } ) ``` ### Translation Translate call transcripts for international teams: ```python config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels=True, speech_understanding={ "request": { "translation": { "target_languages": ["en"], # Translate to English "match_original_utterance": True, # Per-utterance translations "formal": True } } } ) ``` ### Redact PII Text and Audio ```python config = aai.TranscriptionConfig( redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.us_social_security_number, PIIRedactionPolicy.account_number, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True, # Generate de-identified audio ) # After transcription print(transcript.text) # PII redacted in text print(transcript.redacted_audio_url) # PII bleeped in audio ``` ## How Do I Process the Response from the API? ### Processing Pre-recorded Responses ```python expandable def process_call_transcript(transcript): """ Extract and process all relevant data from a pre-recorded call transcript """ call_data = { "id": transcript.id, "duration": transcript.audio_duration, # Already in seconds "confidence": transcript.confidence, "full_text": transcript.text, } # Process speaker utterances speakers = {} for utterance in transcript.utterances: speaker = utterance.speaker if speaker not in speakers: speakers[speaker] = { "utterances": [], "total_speaking_time": 0, "word_count": 0 } speakers[speaker]["utterances"].append({ "text": utterance.text, "start": utterance.start, "end": utterance.end, }) speakers[speaker]["total_speaking_time"] += (utterance.end - utterance.start) / 1000 speakers[speaker]["word_count"] += len(utterance.text.split()) call_data["speakers"] = speakers # Extract summary if transcript.summary: call_data["summary"] = transcript.summary # Analyze sentiment if transcript.sentiment_analysis_results: sentiments = [r.sentiment for r in transcript.sentiment_analysis_results] call_data["sentiment_breakdown"] = { "positive": sentiments.count("POSITIVE"), "neutral": sentiments.count("NEUTRAL"), "negative": sentiments.count("NEGATIVE"), } # Calculate statistics total_duration = transcript.audio_duration call_data["statistics"] = { "total_speakers": len(speakers), "total_words": sum(s["word_count"] for s in speakers.values()), "speaking_distribution": { speaker: { "percentage": (data["total_speaking_time"] / total_duration) * 100, "minutes": data["total_speaking_time"] / 60, } for speaker, data in speakers.items() }, } return call_data result = process_call_transcript(transcript) print(f"Call had {result['statistics']['total_speakers']} speakers") print(f"Sentiment: {result.get('sentiment_breakdown', {})}") ``` ## Additional Resources - [Universal-3.5 Pro Pre-recorded Documentation](/pre-recorded-audio) - [Universal-3.5 Pro Streaming Documentation](/streaming/getting-started/transcribe-streaming-audio) - [Speaker Diarization Guide](/pre-recorded-audio/label-speakers) - [Multichannel Streaming Guide](/streaming/label-speakers-and-separate-channels) - [Speaker Identification Guide](/speech-understanding/speaker-identification) - [Translation Guide](/speech-understanding/translation) - [PII Redaction Guide](/guardrails/redact-pii-from-transcripts) - [Getting Started Guide](/pre-recorded-audio/getting-started/transcribe-an-audio-file) - [API Playground](https://www.assemblyai.com/playground/streaming) - [Changelog](https://www.assemblyai.com/changelog) - [Support](https://www.assemblyai.com/contact/support) --- # Integrations URL: https://www.assemblyai.com/docs/integrations Source: docs/integrations.mdx Navigation: Overview > Use cases & integrations > Integrations Description: Integrations documentation. AssemblyAI seamlessly integrates with a variety of tools and platforms to enhance your workflow. Whether you're looking for no-code solutions or developer tools, we've got you covered. ## Voice agent orchestrators } href="/voice-agents/livekit-universal-3-5-pro" > Use AssemblyAI with Livekit's voice agent orchestrator. } href="/voice-agents/pipecat-universal-3-5-pro" > Use AssemblyAI with Pipecat's voice agent orchestrator. ## No-code tools Connect AssemblyAI with 5000+ apps using Zapier's automation platform. Build complex automation scenarios with Make's visual workflow builder. Automate workflows with n8n's automation platform and AssemblyAI. Test and explore AssemblyAI's APIs using Postman collections. ## Meeting transcriber tools Setup a Zoom meeting bot with AssemblyAI and Recall.ai. Integrate AssemblyAI with Zoom RTMS for meeting transcription. ## Telephony tools Build voice agents with Telnyx and AssemblyAI using Pipecat or LiveKit. Process audio from Twilio calls and voice messages. Setup a transcription pipeline for Amazon Connect recordings with AssemblyAI. Setup a transcription pipeline for Genesys Cloud to AssemblyAI. ## Community-maintained tools Integrate AssemblyAI with LangChain for advanced language model applications. } href="/integrations/vercel-ai-sdk" > Transcribe audio with the AI SDK's unified API using the @ai-sdk/assemblyai provider. Create automated workflows with Microsoft Power Automate to process audio and extract insights. Use AssemblyAI with Microsoft's Semantic Kernel framework. Create automated workflows with this open-source alternative to Zapier. Build production-ready NLP applications with Haystack integration. Transcribe audio with Cloudflare Workers and AssemblyAI. Use AssemblyAI with Relay.app's workflow automation platform. Add speech-to-text capabilities to your Bubble.io no-code applications. Build event-driven workflows with Pipedream's integration platform. Add speech-to-text transcription to your Drupal site through the AI module. Looking to integrate with a different tool? Check out our [API Reference](/api-reference) or [contact our team](https://www.assemblyai.com/contact) for custom integration support. --- # Universal 3.5 Pro Realtime on LiveKit URL: https://www.assemblyai.com/docs/voice-agents/livekit-universal-3-5-pro Source: docs/voice-agents/livekit-universal-3-5-pro.mdx Navigation: Overview > Use cases & integrations > Integrations > Voice agent orchestrators Description: Integrate AssemblyAI's Universal 3.5 Pro Realtime speech-to-text model into a LiveKit voice agent ## Overview This guide covers integrating AssemblyAI's **Universal 3.5 Pro Realtime** speech-to-text model into a LiveKit voice agent using the [Agents framework](https://docs.livekit.io/agents/). **Universal 3.5 Pro Realtime is our flagship next-generation streaming model for voice agents** — multilingual and promptable, with [conversation context](#conversation-context) and [voice focus](#voice-focus). Available on **`livekit-agents` 1.6+** — set `model="universal-3-5-pro"`. AssemblyAI provides the speech-to-text and turn detection in your LiveKit pipeline: ```mermaid flowchart LR U["User audio"] --> STT["AssemblyAI STT
Universal 3.5 Pro Realtime"] STT --> TD["Turn detection"] TD --> LLM["LLM"] LLM --> TTS["TTS"] TTS --> U ``` Once you have an agent running, tune it for what matters most to your use case: Decide when the user is done speaking — modes, defaults, and entity tuning. Shorten the gap between the user finishing and the agent replying. Prompting, key terms, conversation context, and noise handling. Natural barge-in without false triggers from backchannels. For a standalone voice agent without LiveKit or an external LLM, see the [AssemblyAI Voice Agent API](/voice-agents/voice-agent-api), which handles STT, LLM routing, and TTS in a single WebSocket connection. ## Quickstart Get a working, talking agent in a few minutes, then optimize from there. Install the plugin and supporting packages (silero, codecs, dotenv) from PyPI: ```bash pip install "livekit-agents[assemblyai,silero,codecs]~=1.6" \ python-dotenv ``` If you plan to use LiveKit turn detection with `MultilingualModel()`, also install the turn detector plugin: ```bash pip install "livekit-plugins-turn-detector~=1.0" ``` Install the latest `livekit-agents` from PyPI. **Universal 3.5 Pro Realtime**, `agent_context`, and **Voice Focus** require a recent **1.6** release. Automatic conversation-context forwarding and `language_codes` require **1.6.6+**. Older versions won't recognize the `universal-3-5-pro` model and return a validation error. Set your API keys in a `.env` file: ```env LIVEKIT_URL=wss://your-project.livekit.cloud LIVEKIT_API_KEY=your_livekit_api_key LIVEKIT_API_SECRET=your_livekit_api_secret ASSEMBLYAI_API_KEY=your_assemblyai_key # Add API keys for your chosen LLM and TTS providers ``` You can obtain an AssemblyAI API key by [signing up for a free account](https://www.assemblyai.com/dashboard/signup) and navigating to the [API Keys tab](https://www.assemblyai.com/dashboard/home) of the dashboard. The following example uses `turn_detection="stt"` (recommended). Pay close attention to the comments for using with `MultilingualModel()`. ```python expandable from dotenv import load_dotenv from livekit import agents from livekit.agents import AgentSession, Agent, TurnHandlingOptions from livekit.plugins import ( assemblyai, silero, ) # For MultilingualModel, uncomment the following: # from livekit.plugins.turn_detector.multilingual import MultilingualModel load_dotenv() class Assistant(Agent): def __init__(self) -> None: super().__init__(instructions="You are a helpful voice AI assistant.") async def entrypoint(ctx: agents.JobContext): await ctx.connect() session = AgentSession( stt=assemblyai.STT( model="universal-3-5-pro", min_turn_silence=100, max_turn_silence=1000, # When turn_detection="stt", override plugin default of 100. # If using MultilingualModel(), the plugin defaults (min: 100, max: 100) work well. Omit min_turn_silence and max_turn_silence above if preferred. vad_threshold=0.3, # Match Silero's activation_threshold # continuous_partials is True by default in the LiveKit plugin — steady ~3s partials during long turns. # interruption_delay=0, # Optional: faster first partial (~300ms effective). Default: 500 (~800ms effective). ), # llm=your_llm_plugin(), # Add your LLM provider here # tts=your_tts_plugin(), # Add your TTS provider here vad=silero.VAD.load( activation_threshold=0.3, # Match AssemblyAI's internal VAD threshold ), turn_handling=TurnHandlingOptions( turn_detection="stt", # To use LiveKit's turn detection instead, replace the line above with: # turn_detection=MultilingualModel(), endpointing={"min_delay": 0}, # Avoid additive delay in STT mode # If using MultilingualModel(), set these instead: # endpointing={"min_delay": 0.5, "max_delay": 3.0}, ), ) await session.start( room=ctx.room, agent=Assistant(), ) await session.generate_reply( instructions="Greet the user and offer your assistance." ) if __name__ == "__main__": agents.cli.run_app(agents.WorkerOptions(entrypoint_fnc=entrypoint)) ``` For a complete voice agent, you will also need to install LLM and TTS plugins for your chosen providers. See the [LiveKit plugins documentation](https://docs.livekit.io/agents/plugins/) for available options. Start your agent in development mode: ```bash python your_agent_file.py dev ``` Then test it in the LiveKit Playground: 1. Go to [agents-playground.livekit.io](https://agents-playground.livekit.io) 2. Connect to your LiveKit Cloud project (same credentials as your `.env`) 3. Click **Connect**: a room will be created, your agent will join, and you can start talking ## Parameters reference ### **Universal 3.5 Pro Realtime** parameters These are the key parameters to tune for LiveKit when using **Universal 3.5 Pro Realtime**: The streaming model. Defaults to `"universal-3-5-pro"`, our recommended flagship model. Accuracy/latency preset: `"min_latency"`, `"balanced"`, or `"max_accuracy"`. Sets the defaults for mode-dependent fields (`min_turn_silence`, `vad_threshold`, `interruption_delay`, `continuous_partials`, `previous_context_n_turns`); any value you set explicitly still takes precedence. Leave unset for the server's default preset. Connect-time only — cannot be changed with `update_options`. Universal-3.5 Pro only. See [Optimizing accuracy and latency](/streaming/getting-started/optimizing-accuracy-and-latency). List of terms to boost recognition for. Applied automatically alongside any contextual prompt. Contextual prompt — a natural-language description of what the audio is about (domain, scenario, or full details). Transcription behavior is built in and optimized automatically. Your agent's most recent spoken reply, up to 1750 characters, used as context for transcribing the next user turn. The 1750-character cap is enforced client-side — explicit values over the cap raise `ValueError` (both at construction and via `update_options`); only automatic forwarding truncates silently. Can be set at construction time and updated mid-stream with `update_options`. Universal-3.5 Pro only. See [Conversation context](#conversation-context). Steer transcription toward one or more expected languages instead of auto-detecting across all supported languages. Accepts a single code (`"es"`) or a list (`["en", "es"]`); entries are normalized to ISO 639-1 (`"en-US"`, `"english"` → `"en"`). Up to 10 codes. Leave unset for the default multilingual behavior. Universal-3.5 Pro only. Can be updated mid-session via `update_options`; pass an empty list to clear steering. Requires `livekit-agents` **1.6.6+**. How many prior conversation entries are carried forward automatically. Range `0`–`100`; `0` disables carryover. Leave unset for the server default (5). Connect-time only — cannot be changed with `update_options`. Universal-3.5 Pro only. Milliseconds of silence before a speculative end-of-turn check. When the check fires, the model looks for terminal punctuation to decide whether the turn has ended. Maximum milliseconds of silence before the turn is forced to end, regardless of punctuation. **The LiveKit plugin defaults to `100`.** Set to `1000` when using `turn_detection="stt"`. AssemblyAI's internal Silero VAD threshold. **Universal 3.5 Pro Realtime** defaults to `0.3`, unlike Universal-Streaming's `0.4`. Align with LiveKit's Silero `activation_threshold` for consistent behavior. Server-side noise suppression that isolates the primary speaker. `"near-field"` for close-talking mics, `"far-field"` for distant capture. Connect-time only. Universal-3.5 Pro only. See [Voice focus](#voice-focus). How aggressively `voice_focus` suppresses background audio. `0.0`–`1.0`; higher is more aggressive. Connect-time only. Universal-3.5 Pro only. Whether to emit additional partial transcripts during long turns at a steady ~3 second cadence. When enabled (default on both the API and the LiveKit plugin), additional partials covering the full turn transcript are emitted approximately every 3 seconds while speech continues. When disabled, only one early partial is emitted near turn start. The first partial (at 750ms) is unaffected. Useful when downstream consumers (LLMs, UI, eager inference) need frequent updates during long, uninterrupted turns. See [Continuous partials](/streaming/getting-started/transcribe-streaming-audio) for details. How soon the first partial transcript is emitted during a turn, in milliseconds. Range: `0`–`1000`. Lower values produce faster time to first token (TTFT) for barge-in and speculative inference; higher values produce more confident first partials. The server adds a minimum of 300ms on top of the configured value (`interruption_delay: 0` → ~300ms effective, `interruption_delay: 500` → ~800ms effective). See [Tuning early partial timing](/streaming/getting-started/transcribe-streaming-audio) for details. **Universal 3.5 Pro Realtime** code-switches natively between supported languages. This parameter controls whether `language_code` and `language_confidence` are included in turn messages. Defaults to `true` in the LiveKit plugin, but `false` when using the API directly. ### General STT parameters These parameters apply to all AssemblyAI streaming models and can remain the same between models: The sample rate of the audio stream. The encoding of the audio stream. Allowed values: `pcm_s16le`, `pcm_mulaw`. ### Legacy parameters These parameters apply to the `universal-streaming-english` and `universal-streaming-multilingual` AssemblyAI streaming models, but **do not affect Universal 3.5 Pro Realtime**: Confidence threshold for end-of-turn detection. **Universal 3.5 Pro Realtime** uses punctuation-based turn detection instead. Whether to return formatted final transcripts. **Universal 3.5 Pro Realtime** always returns formatted transcripts, so this parameter no longer applies. ## Turn detection In LiveKit, how your agent detects the end of a user's turn is controlled by the [`turn_detection`](https://docs.livekit.io/reference/agents/turn-handling-options/#turn_detection) parameter inside [`TurnHandlingOptions`](https://docs.livekit.io/reference/agents/turn-handling-options/), which is passed to `AgentSession` via the `turn_handling` argument. **Universal 3.5 Pro Realtime** uses a **punctuation-based turn detection system**, which checks for terminal punctuation (`.` `?` `!`) after periods of silence rather than using a confidence score. This means the `min_turn_silence` and `max_turn_silence` parameters you pass to AssemblyAI directly control when transcripts are emitted and when turns end. For more details on how this works, see [Configuring turn detection](/streaming/getting-started/transcribe-streaming-audio). When not explicitly provided, the default endpointing parameters for **Universal 3.5 Pro Realtime** differ on **LiveKit** versus using **AssemblyAI's API directly**: - LiveKit AssemblyAI plugin defaults: - `min_turn_silence=100` - `max_turn_silence=100` - AssemblyAI API defaults: - `min_turn_silence=100` - `max_turn_silence=1000` However, you can always override these by passing your own preferred values explicitly. Misconfiguring these parameters is the most common cause of poor performance — read the recommended values per mode below. ### Default parameter differences **Universal 3.5 Pro Realtime's** endpointing is controlled by two AssemblyAI API parameters, `min_turn_silence` and `max_turn_silence`, that you pass to the STT plugin. These are separate from LiveKit's `endpointing.min_delay` and `endpointing.max_delay` (set inside `TurnHandlingOptions`). | Parameter | AssemblyAI API default | LiveKit plugin default | Description | | --------------------- | ---------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `min_turn_silence` | `100` ms | `100` ms | Silence before a speculative end-of-turn check. If terminal punctuation (`.` `?` `!`) is found, the turn ends. If not, a partial is emitted and the turn continues. | | `max_turn_silence` | **`1000`** ms | **`100`** ms | Maximum silence before forcing the turn to end, regardless of punctuation. | The LiveKit plugin defaults are optimized for third-party turn detection models, where you want transcripts handed off as fast as possible. When using `turn_detection="stt"`, you should explicitly set `max_turn_silence=1000` if you'd like to mimic the behavior of streaming directly to the API without LiveKit. **Tuning endpointing parameters** **These are the default values used when no parameters are explicitly provided.** You will likely need to experiment with different values depending on your use case: - **Increase `min_turn_silence`**: when brief pauses cause the speculative EOT check to fire too early, ending turns on terminal punctuation before the user has finished speaking. - **Increase `max_turn_silence`**: when the forced turn end is cutting off users mid-thought or splitting entities like phone numbers across turns, a higher value lets the model wait longer before forcing the turn to end when the model is unsure. See the [Entity splitting tradeoff](#entity-splitting-tradeoff) section for examples. ### STT-based Turn Detection (recommended) With `turn_detection="stt"`, AssemblyAI's built-in punctuation-based turn detection determines when the user has finished speaking. AssemblyAI's `end_of_turn` signals are then used directly by LiveKit to commit the turn. In this mode, we recommend explicitly setting `min_turn_silence=100` and `max_turn_silence=1000`. These are AssemblyAI's API defaults and provide a good balance of responsiveness and accuracy. **The LiveKit plugin defaults to `min_turn_silence=100` and `max_turn_silence=100`, which might be too aggressive for STT-based turn detection.** **Recommended starting parameters** (set on `assemblyai.STT()`, not on `AgentSession`): | Parameter | Default | Description | | ------------------ | --------- | -------------------------------------------------------------------- | | `min_turn_silence` | `100` ms | Silence duration before a speculative end-of-turn (EOT) check fires. | | `max_turn_silence` | `1000` ms | Maximum silence before a turn is forced to end. | **How it works:** 1. User speaks → audio streams to AssemblyAI 2. User pauses for `100ms` → AssemblyAI checks for terminal punctuation 3. If terminal punctuation (`.` `?` `!`) → turn ends immediately 4. If no terminal punctuation → partial emitted, turn continues waiting 5. If silence reaches `1000ms` → turn is forced to end regardless of punctuation **Endpointing `min_delay` is additive in STT mode** LiveKit's endpointing `min_delay` (default **0.5 seconds**) is applied **on top of** AssemblyAI's own endpointing. In STT mode, this delay starts *after* the STT end-of-speech signal, meaning it adds up to `500ms` of extra latency by default. **Set `endpointing={"min_delay": 0}`** inside `TurnHandlingOptions` to avoid this. AssemblyAI's own endpointing parameters (`min_turn_silence` and `max_turn_silence`) already control the timing, so an additional delay on the LiveKit side is unnecessary latency. See [Latency](#latency) for the full picture. ```python from livekit.agents import AgentSession, TurnHandlingOptions session = AgentSession( turn_handling=TurnHandlingOptions( turn_detection="stt", endpointing={"min_delay": 0}, # Avoid additive delay in STT mode ), stt=assemblyai.STT( model="universal-3-5-pro", min_turn_silence=100, # Silence (ms) before a speculative end-of-turn check max_turn_silence=1000, # Max silence (ms) before forcing the turn to end vad_threshold=0.3, ), vad=silero.VAD.load( activation_threshold=0.3, ), ) ``` ### LiveKit turn detection (with `MultilingualModel()`) As a third-party turn detection model, LiveKit's [turn detector](https://docs.livekit.io/agents/logic/turns/turn-detector/) runs on top of STT output to make turn decisions. AssemblyAI's role is then just to provide transcripts as quickly as possible, while the turn detection model decides when the user is actually done speaking. Use `MultilingualModel()` rather than `EnglishModel()`, as **Universal 3.5 Pro Realtime** supports 18 languages. `MultilingualModel()` covers support for all of these languages. The LiveKit plugin defaults of `min_turn_silence=100` and `max_turn_silence=100` work well here, as **`max_turn_silence` is brought down to match `min_turn_silence` so that transcripts are handed off to the turn detection model as fast as possible.** **`MultilingualModel` parameters** (set inside `TurnHandlingOptions` → `endpointing`, not on the STT plugin): | Parameter | Default | Description | | ----------------------- | ------- | -------------------------------------------------------------------------------------------------------------------------- | | `endpointing.min_delay` | `0.5` s | Time to wait before committing a turn when the model predicts a likely boundary. | | `endpointing.max_delay` | `3.0` s | Maximum time to wait when the model predicts the user will continue speaking. Has no effect without a turn detector model. | **How it works:** 1. User speaks → audio streams to AssemblyAI 2. User pauses for `100ms` → AssemblyAI emits transcript (final and partial are the same) immediately 3. LiveKit's `MultilingualModel()` evaluates the transcript in conversational context 4. If the model predicts a likely turn boundary → waits `min_delay` (`0.5s`) then commits the turn 5. If the model predicts the user will continue → waits up to `max_delay` (`3.0s`) for more speech ```python from livekit.agents import AgentSession, TurnHandlingOptions from livekit.plugins.turn_detector.multilingual import MultilingualModel session = AgentSession( turn_handling=TurnHandlingOptions( turn_detection=MultilingualModel(), endpointing={ "min_delay": 0.5, # Time (s) to wait before committing a turn when the model is confident "max_delay": 3.0, # Max time (s) to wait when the model is not confident }, ), stt=assemblyai.STT( model="universal-3-5-pro", vad_threshold=0.3, ), vad=silero.VAD.load( activation_threshold=0.3, ), ) ``` ### Other turn detection modes - **`vad`**: - Detect end of turn from speech and silence data alone using [Silero VAD](https://docs.livekit.io/agents/logic/turns/vad/). - Turn boundaries are determined purely by voice activity without semantic context. - AssemblyAI's turn detection parameters still control when transcripts are emitted, but it is recommended to leave them at the plugin defaults (`min_turn_silence=100`, `max_turn_silence=100`) so transcripts arrive as quickly as possible. - **`manual`**: - Disable automatic turn detection entirely. - You control turns explicitly using `session.commit_user_turn()`, `session.clear_user_turn()`, and `session.interrupt()`. - See the [manual turn control docs](https://docs.livekit.io/agents/logic/turns/#manual-turn-control) for details. ### VAD configuration With `turn_detection="stt"`, AssemblyAI also sends `SpeechStarted` events that LiveKit uses for barge-in/interruption handling. `SpeechStarted` is only emitted when the model produces a transcript. Silero VAD is not strictly required in this mode, but it is still recommended as Silero runs locally and it can be faster than waiting for AssemblyAI's `SpeechStarted` signal. LiveKit respects whichever signal arrives first, so Silero provides faster interruption while AssemblyAI's signal serves as a reliable backup. With `MultilingualModel()`, Silero VAD is **required**, as it is the only source of `START_OF_SPEECH` events for interruption in this mode. AssemblyAI's `SpeechStarted` event is not used. #### Threshold alignment LiveKit's Silero VAD defaults to an `activation_threshold` of **0.5**. AssemblyAI's `vad_threshold` defaults to **0.3**. For best performance, we recommend setting both to **0.3**. Both should be adjusted together to the same value to ensure accurate transcription and consistent barge-in thresholds. When the thresholds are mismatched, you get a dead zone: if Silero is at `0.5` and AssemblyAI is at `0.3`, AssemblyAI will be actively transcribing speech that LiveKit hasn't detected yet, delaying interruption. Keeping them aligned eliminates this. ```python session = AgentSession( stt=assemblyai.STT( model="universal-3-5-pro", vad_threshold=0.3, # AssemblyAI's internal VAD onset ), vad=silero.VAD.load( activation_threshold=0.3, # Match AssemblyAI's threshold ), ) ``` If you're in a noisy environment and receiving false speech triggers, raise both `stt.vad_threshold` and `vad.activation_threshold` thresholds together. ### Entity splitting tradeoff Lower `min_turn_silence` and `max_turn_silence` values produce faster transcripts but can split entities or utterances across turns. The two parameters affect this differently. #### `min_turn_silence` too low - Speculative check fires too early, splitting entities on punctuation. - **Example:** _User spells out an email address with brief pauses between parts. The speculative check fires at 100ms of silence, and the model adds terminal punctuation to each segment, ending the turn prematurely._ ```text # With (min_turn_silence=100, max_turn_silence=1000) "It's John." → FINAL (100ms pause, check fires, period found → turn ends) "Smith." → FINAL "At gmail.com." → FINAL # With (min_turn_silence=400, max_turn_silence=1000) "It's john.smith@gmail.com." → FINAL (single turn, properly formatted) ``` #### `max_turn_silence` too low - Forced turn-end cuts off user mid-thought. - **Example:** _User pauses longer than 1 second to think mid-sentence. The forced end fires at 1000ms, splitting the utterance into two turns regardless of punctuation._ ```text # With (min_turn_silence=100, max_turn_silence=1000) "I wanted to check on my order from..." → FINAL (1000ms silence, forced end) "last Tuesday, order number 4829." → FINAL (new turn) # With (min_turn_silence=100, max_turn_silence=2000) "I wanted to check on my order from last Tuesday, order number 4829." → FINAL (single turn) ``` **Universal 3.5 Pro Realtime's** formatting is significantly better when it has full context in a single turn. Email addresses, phone numbers, credit card numbers, and physical addresses all benefit from this. LLMs downstream can usually piece together split entities, but if your use case involves alphanumeric dictation or entity extraction, consider increasing `min_turn_silence` and `max_turn_silence` during those portions of the conversation. You can [update configuration mid-stream](#dynamic-configuration) to raise `max_turn_silence` temporarily (e.g., to `2000`–`4000` ms) when expecting entity input, then lower it again afterward. Even when using third-party turn detection, you may want to increase `min_turn_silence` or `max_turn_silence` if users are likely to speak slowly or dictate entities. While this adds latency, it improves accuracy by giving the model more audio context before emitting a transcript and keeping the full entity complete within the same turn. ## Latency A voice agent feels responsive when the gap between the user finishing and the agent replying is short. Start with the **`mode` preset** — the highest-level dial for the accuracy/latency trade-off. It sets sensible defaults for the fine-grained levers below, so you can pick a target and tune from there: ```python stt=assemblyai.STT( model="universal-3-5-pro", mode="balanced", # "min_latency" (fastest) · "balanced" · "max_accuracy" (best quality) ) ``` `mode` is set at construction time (it can't be changed mid-session) and influences the defaults of `min_turn_silence`, `vad_threshold`, `interruption_delay`, `continuous_partials`, and `previous_context_n_turns`. Any value you set explicitly still wins. Leave it unset to use the server's default preset. See [Optimizing accuracy and latency](/streaming/getting-started/optimizing-accuracy-and-latency). From there, fine-tune the individual levers: - **Endpointing delay (STT mode).** LiveKit's `endpointing.min_delay` is applied *on top of* AssemblyAI's own end-of-turn timing and adds up to `500ms` by default. Set `endpointing={"min_delay": 0}` when using `turn_detection="stt"` — see [STT-based turn detection](#stt-based-turn-detection-recommended). - **End-of-turn timing.** `min_turn_silence` (speculative check) and `max_turn_silence` (forced end) directly control how soon a turn ends. Lower is faster but risks splitting entities — see [Turn detection](#turn-detection). - **Time to first partial.** `interruption_delay` controls how soon the first partial is emitted, which drives faster barge-in and speculative inference. The server adds a minimum of `300ms` on top of the configured value. ```python # Tune first-partial timing for faster barge-in stt.update_options( interruption_delay=0, # ~300ms effective TTFT ) ``` - **Sample rate.** Use 16 kHz (`sample_rate=16000`). Higher rates don't improve accuracy and only add bandwidth. - **Continuous partials.** `continuous_partials` (on by default) emits a partial every ~3 seconds during long turns. Leave it on for steady mid-turn updates, or disable it if you only need a single early partial. - **Skip client-side preprocessing.** Don't run your own noise cancellation before audio reaches the model — the artifacts it introduces usually hurt accuracy more than the original noise. Use server-side [Voice Focus](#voice-focus) instead. ### Latency breakdown | Stage | Typical | Controlled by | | ------------------------------------------ | ---------------------------------------------------- | ----------------------- | | Network round trip | ~50 ms | — | | Speech-to-text | ~200–300 ms | model | | First partial (TTFT) | configured `interruption_delay` + ~300 ms server min | `interruption_delay` | | End of turn (terminal punctuation found) | `min_turn_silence` (default 100 ms) | `min_turn_silence` | | End of turn (no punctuation, forced) | up to `max_turn_silence` | `max_turn_silence` | | LiveKit endpointing (STT mode) | + `min_delay` (set to `0`) | `endpointing.min_delay` | ## Accuracy **Universal 3.5 Pro Realtime** is accurate out of the box. When you need more — domain vocabulary, proper nouns, noisy audio — reach for these levers. For entity-heavy dictation, also tune turn detection (see [Entity splitting tradeoff](#entity-splitting-tradeoff)), and note that the high-level [`mode` preset](#latency) shifts the overall accuracy/latency balance (use `max_accuracy` to favor quality). ### Prompting **Beta feature** Prompting is considered a beta feature for **Universal 3.5 Pro Realtime**. While it can be a powerful tool for improving accuracy in certain use cases, **we recommend starting without a `prompt` to first establish baseline performance.** Once the baseline has been tested, you can add context to further optimize for your use case (e.g., language mix to expect (e.g., English and Hindi), use case or domain (e.g., medical, legal), etc.). **Universal 3.5 Pro Realtime** supports a `prompt` parameter for [contextual prompting](/streaming/prompting-and-keyterms) — a description of what the audio is about. Transcription behavior (verbatim output, punctuation, turn detection) is built in and optimized automatically; the prompt carries context, not instructions. ```python stt=assemblyai.STT( model="universal-3-5-pro", prompt="Customer support call about an internet service outage.", ) ``` **Tips:** - **Start with no prompt**: Universal 3.5 Pro Realtime delivers strong accuracy out of the box — only add context if domain-specific vocabulary is being misrecognized. - **Describe the conversation**: domain, scenario, or full details — start broad, and add only details your application actually knows (see the [three context levels](/streaming/prompting-and-keyterms#contextual-prompting)). - **Include known names and identifiers**: caller name, account or order IDs, products — detailed context helps the model spell them correctly. ### Key terms Instead of `prompt`, use `keyterms_prompt` to boost recognition of specific names, brands, or domain terms: ```python stt=assemblyai.STT( model="universal-3-5-pro", keyterms_prompt=["AssemblyAI", "LiveKit", "Universal 3.5 Pro Realtime"], ) ``` ### Language selection By default, **Universal 3.5 Pro Realtime** auto-detects and code-switches across all supported languages. If you know which languages a caller will speak, `language_codes` steers transcription toward them for better accuracy on short or ambiguous turns. It can also be updated mid-session (e.g. when a caller switches languages) via `update_options`. Requires `livekit-agents` **1.6.6+**. ```python stt=assemblyai.STT( model="universal-3-5-pro", language_codes=["en", "es"], # single code or list, up to 10; ISO 639-1 ) ``` ### Conversation context Give the model both sides of the dialog so it transcribes the next user turn more accurately. **Universal 3.5 Pro Realtime** keeps a short, per-session memory of the conversation from two sources: - **The agent half** — what your agent just said, forwarded automatically into `agent_context` by `AgentSession` on `livekit-agents` 1.6.6+ (no wiring needed). - **The user half** — prior STT-finalized user turns, carried forward automatically (no configuration needed). With the agent's question in context, the model can anticipate the answer, sharpen entity recognition, and disambiguate similar-sounding words. For example, after your agent asks `"What's your email address?"`, `agent_context` lets the model produce `"user@assemblyai.com"` instead of `"user at assemblyai dot com"`. This has the biggest impact on short replies (`"yes"`, `"7pm"`, single names) and spelled-out entities. See [Conversation context](/streaming/universal-3-5-pro/context-carryover) for the full reference. | Parameter | Type | Description | | -------------------------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `agent_context` | string | Your agent's most recent spoken reply (what your TTS just said), up to 1750 characters. Set at construction time to seed an opening greeting; updated automatically after each agent reply on 1.6.6+. | | `previous_context_n_turns` | integer | How many prior conversation entries are carried forward automatically. Connect-time only. Range `0`–`100`; `0` disables carryover; leave unset for the server default (5). | `agent_context` and `previous_context_n_turns` are supported only on Universal-3.5 Pro (`universal-3-5-pro`). `previous_context_n_turns` is set at construction time and cannot be changed with `update_options()`. #### Automatic on 1.6.6+ On `livekit-agents` **1.6.6+**, `AgentSession` forwards each assistant reply into `agent_context` automatically — no wiring needed. Replies longer than the 1750-character limit are truncated automatically, keeping the tail (the end of the reply — usually the question — has the most biasing value). To opt out, use the session-level toggle: ```python session = AgentSession( stt=assemblyai.STT(model="universal-3-5-pro"), stt_context_options={"forward_chat_context": False}, # default: True ) ``` While automatic forwarding is on, any `agent_context` you set manually is replaced by the next assistant reply. To manage `agent_context` yourself (e.g. custom trimming or filtering), set `forward_chat_context: False` and push it via `update_options(agent_context=...)` after each reply. To seed the model with your agent's opening greeting, pass `agent_context` on `assemblyai.STT(...)` at construction time — automatic forwarding takes over from the first reply onward. If you're on a version of `livekit-agents` that predates **1.6.6**, or you want to manage `agent_context` yourself, wire it via the `conversation_item_added` event. LiveKit emits it for every conversation item; filter for the assistant's messages and forward the (trimmed) text to the STT plugin via `update_options(agent_context=...)`. No reconnect is required. A complete runnable example is on GitHub: [including-agent-context](https://github.com/dlange-aai/assemblyai-livekit-examples/tree/main/including-agent-context). ```python from livekit.agents import AgentSession, ConversationItemAddedEvent from livekit.plugins import assemblyai AGENT_CONTEXT_MAX_CHARS = 1750 session = AgentSession( stt=assemblyai.STT( model="universal-3-5-pro", ), stt_context_options={"forward_chat_context": False}, # Disable automatic forwarding on 1.6.6+ # llm=your_llm_plugin(), # tts=your_tts_plugin(), ) @session.on("conversation_item_added") def _on_conversation_item_added(ev: ConversationItemAddedEvent) -> None: # Only forward the agent's own replies as context for the next user turn. if ev.item.type != "message" or ev.item.role != "assistant": return agent_stt = session.stt if not isinstance(agent_stt, assemblyai.STT): return spoken = ev.item.text_content if not spoken: return # Push the latest reply; earlier turns are carried forward automatically. agent_stt.update_options(agent_context=spoken[-AGENT_CONTEXT_MAX_CHARS:]) ``` #### LiveKit Inference LiveKit Inference (`stt="assemblyai/universal-3-5-pro"`) receives the same automatic conversation-context forwarding by default, with the same `stt_context_options={"forward_chat_context": False}` opt-out. `language_codes` is plugin-only for now. ### Voice focus Voice Focus isolates the primary speaker and suppresses background noise — chatter, keyboard clicks, fan hum, room echo — **server-side, before audio reaches the model**. Use it instead of client-side noise cancellation, which tends to introduce artifacts that hurt accuracy more than the noise itself. | Parameter | Type | Description | | ----------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------ | | `voice_focus` | string | `"near-field"` for headsets, handsets, and other close-talking mics; `"far-field"` for conference rooms, laptop mics, and other distant capture. | | `voice_focus_threshold` | float | Optional. `0.0`–`1.0`; higher values suppress background audio more aggressively. | Both are connection-time parameters on the Universal-3.5 Pro and cannot be changed with `update_options()`. See [Voice Focus](/streaming/voice-focus) for details. ```python stt=assemblyai.STT( model="universal-3-5-pro", voice_focus="far-field", # "near-field" for close-talking mics voice_focus_threshold=0.5, # Optional: 0.0–1.0, higher = more aggressive ) ``` ## Interruption handling In voice agent conversations, users often produce backchannel utterances ("mhm", "yeah", "um", "okay") while the agent is speaking. These short fillers can trigger LiveKit's interruption logic, causing the agent to stop mid-sentence even though the user didn't intend to interrupt. LiveKit Cloud users should reach for adaptive interruption handling first. Self-hosted deployments can combine the two custom filters described below. ### Recommended: Adaptive interruption handling LiveKit Agents v1.5.0 introduced [adaptive interruption handling](https://docs.livekit.io/agents/logic/turns/adaptive-interruption-handling/), a server-side model that classifies overlapping speech as a real interrupt or a backchannel. On a false interrupt the agent's TTS resumes from where it left off. No re-generation is needed. The feature is recommended for LiveKit Cloud users and is included on all Cloud plans. ```python session = AgentSession( stt=assemblyai.STT(model="universal-3-5-pro"), # llm=your_llm_plugin(), # tts=your_tts_plugin(), vad=silero.VAD.load(), interruption={"mode": "adaptive"}, # default in v1.5.0+ ) ``` See LiveKit's [adaptive interruption handling docs](https://docs.livekit.io/agents/logic/turns/adaptive-interruption-handling/) for full requirements and behavior. ### Self-hosted alternative If you're self-hosting LiveKit, can't use adaptive interruption handling, or need explicit control over which utterances count as filler, the two complementary filters below provide strong guardrails for interruption and barge-in handling. They have non-overlapping failure modes: the backchannel filter stops known fillers before they enter the pipeline, while the buffer-clearing filter catches any short utterance (including unknown fillers or stutters) that slips past. Running both provides the strongest coverage. A working reference implementation is available on [GitHub](https://github.com/AssemblyAI-Solutions/livekit-interruption-filters). The backchannel filter intercepts STT events at the `stt_node` level, before they reach LiveKit's `audio_recognition`. It checks each transcript event for known disfluencies and backchannels ("um", "mhm", "yeah", "okay", etc.) and drops the event entirely. Because the event never enters the pipeline, none of the downstream orchestration (interrupt gates, end-of-turn detection, and preemptive LLM generation) ever reacts to it. The filter is implemented as a mixin class that wraps `Agent.stt_node`: ```python from __future__ import annotations import logging import string import time from collections.abc import AsyncIterable from livekit import rtc from livekit.agents import Agent, stt from livekit.agents.voice import ModelSettings # "yes" / "no" deliberately omitted - in a booking flow a bare "yes" # is a real confirmation. Edit for your domain. BACKCHANNELS = frozenset({ "mhm", "mm", "mmhm", "mmhmm", "uh", "uhhuh", "huh", "um", "umm", "uhm", "er", "erm", "hmm", "hm", "ah", "oh", "yeah", "yep", "yup", "okay", "ok", "right", "alright", "gotcha", }) _TRANSCRIPT_TYPES = { stt.SpeechEventType.INTERIM_TRANSCRIPT, stt.SpeechEventType.PREFLIGHT_TRANSCRIPT, stt.SpeechEventType.FINAL_TRANSCRIPT, } _PUNCT_STRIP = str.maketrans("", "", string.punctuation) log = logging.getLogger("backchannel_stt_filter") def _is_all_backchannel(text: str) -> bool: """Return True only when every token is a known backchannel.""" tokens = text.lower().translate(_PUNCT_STRIP).split() return bool(tokens) and all(tok in BACKCHANNELS for tok in tokens) class BackchannelSTTFilterMixin: """Drop backchannel-only transcripts while the agent is speaking.""" _FILTER_GRACE_S: float = 1.0 _last_speaking_at: float = 0.0 async def stt_node( self, audio: AsyncIterable[rtc.AudioFrame], model_settings: ModelSettings ): async for ev in Agent.default.stt_node(self, audio, model_settings): if self._should_drop(ev): text = ev.alternatives[0].text if ev.alternatives else "" log.info( "event_filtered transcript=%r ev_type=%s agent_state=%s", text, ev.type, self.session.agent_state, ) continue yield ev def _should_drop(self, ev: stt.SpeechEvent) -> bool: now = time.monotonic() if self.session.agent_state == "speaking": self._last_speaking_at = now elif now - self._last_speaking_at > self._FILTER_GRACE_S: return False if ev.type not in _TRANSCRIPT_TYPES: return False text = ev.alternatives[0].text if ev.alternatives else "" return _is_all_backchannel(text) ``` **How it works:** 1. While the agent is speaking (plus a 1-second grace window after speech ends), the filter inspects each STT transcript event 2. It strips punctuation and checks whether every token in the transcript matches the `BACKCHANNELS` set 3. Pure-filler transcripts like "mhm" or "yeah okay" are dropped. They never reach LiveKit's pipeline 4. Utterances with any non-filler token (e.g., "yeah I want the suite") always pass through The `BACKCHANNELS` set is domain-customizable. "yes" and "no" are deliberately excluded because they often represent genuine confirmations in booking or IVR scenarios. Edit the set for your use case. LiveKit accumulates committed `FINAL` transcripts in a private buffer (`_audio_transcript`) across the user's uncommitted turn. The `_interrupt_by_audio_activity` method checks the running word count against `min_words` to decide whether to pause TTS. Without intervention, two consecutive short fillers like "yeah" + "um" sum to two words and trip the interrupt gate, even though each utterance on its own is below threshold. This filter listens on the `user_input_transcribed` event and wipes the buffer whenever the agent is speaking and the user input falls below the configured `min_words` threshold. Each short utterance is evaluated independently rather than against the accumulated total. This filter requires `interruption.min_words` to be set to `2` or higher. Without it, the word-count gate is disabled and the filter has no effect. ```python from __future__ import annotations import logging from livekit.agents import AgentSession log = logging.getLogger("short_utterance_buffer_filter") def install_short_utterance_filter(session: AgentSession) -> None: """Clear transcript buffers when a short utterance arrives during agent speech.""" @session.on("user_input_transcribed") def _on_user_input_transcribed(ev) -> None: word_count = len(ev.transcript.split()) min_words = session.options.interruption["min_words"] if session.agent_state != "speaking": return if word_count >= min_words: return activity = getattr(session, "_activity", None) recognition = getattr(activity, "_audio_recognition", None) if activity else None if recognition is None: return # Wipe all three transcript buffers so short utterances # don't accumulate past the interrupt threshold. recognition._audio_transcript = "" recognition._audio_interim_transcript = "" recognition._audio_preflight_transcript = "" # Best-effort: abort any in-flight preemptive LLM call # triggered by this short utterance. cancel = getattr(activity, "_cancel_preemptive_generation", None) if callable(cancel): try: cancel() except Exception: log.debug("_cancel_preemptive_generation failed", exc_info=True) log.info( "buffer_cleared transcript=%r words=%d is_final=%s", ev.transcript, word_count, ev.is_final, ) ``` This filter accesses LiveKit private APIs (`_audio_transcript`, `_audio_interim_transcript`, `_audio_preflight_transcript`, `_cancel_preemptive_generation`). Pin your `livekit-agents` version to avoid breakage on minor updates. Tested with `livekit-agents>=1.5`. Three changes are needed: **1. Mix in the backchannel filter on your agent class.** Place it before `Agent` in the class bases so its `stt_node` runs first: ```python from filters.backchannel_stt import BackchannelSTTFilterMixin class MyAgent(BackchannelSTTFilterMixin, Agent): def __init__(self) -> None: super().__init__(instructions="You are a helpful voice AI assistant.") ``` **2. Configure the session with `interruption.min_words >= 2`.** This enables the word-count gate that the buffer-clearing filter depends on: ```python session = AgentSession( stt=assemblyai.STT( model="universal-3-5-pro", min_turn_silence=100, max_turn_silence=1000, vad_threshold=0.3, ), # llm=your_llm_plugin(), # tts=your_tts_plugin(), vad=None, # Recommended: disable VAD so only STT drives interruption turn_handling={ "turn_detection": "stt", "endpointing": {"min_delay": 1.0, "max_delay": 4.0}, "interruption": { "enabled": True, "resume_false_interruption": True, "false_interruption_timeout": 1.5, "min_words": 2, # Required for the buffer-clearing filter }, }, ) ``` **3. Install the buffer-clearing filter on the session**, then start it with your agent: ```python from filters.short_utterance_buffer import install_short_utterance_filter install_short_utterance_filter(session) await session.start(room=ctx.room, agent=MyAgent()) ``` Setting `vad=None` with `turn_detection="stt"` is the recommended setup for both filters. This ensures only STT-based signals drive interruption, avoiding timing races from a competing VAD interrupt path. If you need VAD for faster barge-in, both filters still work: set `interruption.min_words` to `2` and ensure Silero's `activation_threshold` matches `vad_threshold`. **Expected behavior** with both filters active and `min_words=2`: | User utterance (during agent speech) | Backchannel filter | Buffer-clearing filter | Result | | ------------------------------------ | ------------------ | ---------------------- | ------ | | "mhm" | Dropped | - | Agent continues | | "um" | Dropped | - | Agent continues | | "yeah yeah" | Dropped | - | Agent continues | | Unknown short word | Passes through | Buffer cleared | Agent continues | | "yeah I'd like the suite" | Passes through | Passes through | Agent interrupts | | "suite please" | Passes through | Passes through | Agent interrupts | | "mhm" (1+ second after agent stops) | Passes through | - | Agent responds | ## Dynamic configuration You can update `prompt`, `keyterms_prompt`, `agent_context`, `language_codes`, `min_turn_silence`, `max_turn_silence`, `continuous_partials`, and `interruption_delay` during an active session using `update_options` — no reconnect required. This lets you adapt to the conversation stage: boost names while collecting caller details, widen silence windows while a user dictates an email, then tighten them again. ```python # Update one or more options mid-stream stt.update_options( max_turn_silence=3000, # Increase for entity dictation ) # Later, reset to default stt.update_options( max_turn_silence=1000, ) ``` | Conversation stage | Adjustment | | ------------------------------------------- | --------------------------------------------------------------------------------------- | | Caller identification (names, account IDs) | Boost terms with `update_options(keyterms_prompt=[...])` | | Entity dictation (email, phone, address) | Raise `max_turn_silence` to ~`2000`–`4000` ms, then lower it again afterward | | After each agent reply | Handled automatically on 1.6.6+ (see [Conversation context](#conversation-context)) | | Caller switches language | `update_options(language_codes=["es"])` | | Faster barge-in | Lower `interruption_delay` (see [Latency](#latency)) | For more information, see [Updating configuration mid-stream](/streaming/updating-configuration-mid-stream). ## Troubleshooting | Issue | Cause | Solution | | ----------------------------------------- | --------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- | | Extra latency with `turn_detection="stt"` | LiveKit's endpointing `min_delay` is additive in STT mode | Set `endpointing={"min_delay": 0}` inside `TurnHandlingOptions` on `AgentSession` | | No interruption handling | Missing VAD | Ensure `vad=silero.VAD` is set, with `activation_threshold` equal to `vad_threshold` (default `0.3`) | | Turn over-segmentation | `min_turn_silence` too low | Increase from `100` to `200`–`500` | | Entities split across turns | `max_turn_silence` too low | Increase `max_turn_silence` (e.g., `1500`–`3500`) | | Latency on non-terminal utterances | `max_turn_silence` too high | Lower `max_turn_silence` | | Mis-heard names, brands, or jargon | No vocabulary hints | Add `keyterms_prompt`, or supply `prompt`/`agent_context` for context | | Poor accuracy in noisy audio | Background noise or room echo | Enable `voice_focus` (`near-field` or `far-field`) | ## Migrating from another STT provider To balance accuracy, latency, turn-taking, and interruption handling, map your current setup to AssemblyAI using the questions below. Each answer points to the settings that reproduce — and usually improve on — your current behavior. ### How are you detecting end-of-turn today? | Today | Recommended on AssemblyAI | | --- | --- | | Your STT provider's own end-of-turn model (e.g. Deepgram endpointing) | `turn_detection="stt"` with `min_turn_silence=100`, `max_turn_silence=1000`, `endpointing={"min_delay": 0}`. AssemblyAI's punctuation-based end-of-turn replaces it. | | Silence / VAD only | Prefer `turn_detection="stt"` for semantic endpointing, or `turn_detection="vad"` to stay silence-only. Align Silero `activation_threshold` and `vad_threshold` at `0.3`. | | LiveKit's turn-detector model | `turn_detection=MultilingualModel()`, keep the plugin defaults (`min_turn_silence=100`, `max_turn_silence=100`), set `endpointing={"min_delay": 0.5, "max_delay": 3.0}`, and keep Silero VAD (required). | ### Which model and settings are you migrating from? | What you pass today | AssemblyAI equivalent | | --- | --- | | Current model (Deepgram, ElevenLabs, etc.) | `model="universal-3-5-pro"` (recommended flagship) | | Overall accuracy/latency tuning | `mode="min_latency"` / `"balanced"` / `"max_accuracy"` — a one-line starting point before fine-tuning | | Endpointing / silence thresholds | `min_turn_silence` (speculative end-of-turn) and `max_turn_silence` (forced end) — see [Turn detection](#turn-detection) | | Custom vocabulary / keywords | `keyterms_prompt=[...]`; broader domain context → `prompt` | | Formatting / punctuation toggles | On by default — formatted transcripts always (`format_turns` does not apply) | | Language(s) | 18 languages with code-switching; pin one or more languages via `language_codes`; `language_detection` adds language codes | ### VAD, interruptions, and infrastructure | Today | Recommended | | --- | --- | | Silero VAD | Keep it; set `activation_threshold=0.3` to match `vad_threshold` | | Adaptive vs. VAD-based barge-in | LiveKit Cloud → [adaptive interruption handling](#recommended-adaptive-interruption-handling); self-hosted → the [self-hosted filters](#self-hosted-alternative) | | Telephony / SIP routing | `sample_rate=8000` and `encoding="pcm_mulaw"` for 8 kHz telephony | | Client-side noise cancellation | Drop it; use server-side [Voice Focus](#voice-focus) instead | | LiveKit Cloud vs. self-hosting | Cloud unlocks adaptive interruption; self-hosting uses the filters and version pinning | ### Parameter cheat sheet If you're coming from AssemblyAI's standard streaming model: | Change | From | To | | ---------------------------------- | ------------------------------------------ | ---------------------------------------------------------------------------------------- | | Model | `assemblyai.STT()` | `assemblyai.STT(model="universal-3-5-pro")` | | Turn detection | `turn_detection="stt"` or `EnglishModel()` | `turn_detection="stt"` or `MultilingualModel()` | | VAD | Optional | Set `vad=silero.VAD.load()` to match `vad_threshold` | | `min_turn_silence` | `400` (old default) | `100` (new default) | | `max_turn_silence` | `1280` (old default) | `1000` (API default) or `100` (with 3rd-party turn detector) | | `end_of_turn_confidence_threshold` | Configurable | **Not applicable**: **Universal 3.5 Pro Realtime** uses punctuation-based turn detection | | `endpointing.min_delay` (formerly `min_endpointing_delay`) | Default `0.5` | Set to `0` inside `TurnHandlingOptions` when using `turn_detection="stt"` | Migrating a production deployment? [Talk to our team](https://www.assemblyai.com/contact/sales). --- # Universal 3.5 Pro Realtime on Pipecat URL: https://www.assemblyai.com/docs/voice-agents/pipecat-universal-3-5-pro Source: docs/voice-agents/pipecat-universal-3-5-pro.mdx Navigation: Overview > Use cases & integrations > Integrations > Voice agent orchestrators Description: Integrate AssemblyAI's Universal 3.5 Pro Realtime speech-to-text model into a Pipecat voice agent ## Overview This guide covers integrating AssemblyAI's **Universal 3.5 Pro Realtime** speech-to-text model into a [Pipecat](https://docs.pipecat.ai/) voice agent. Everything here applies equally to **Universal-3.5 Pro Streaming** (`universal-3-5-pro`) — both belong to the same U3 Pro family and share every parameter in this guide, so you can swap the `model` string without changing anything else. **Universal 3.5 Pro Realtime is our flagship next-generation streaming model for voice agents** — multilingual and promptable, with [conversation context](#conversation-context) and [voice focus](#voice-focus). Available on **Pipecat 1.4.0+** — set `model="universal-3-5-pro"`. AssemblyAI provides the speech-to-text and (optionally) the turn detection in your Pipecat pipeline: ```mermaid flowchart LR U["User audio"] --> STT["AssemblyAI STT
Universal 3.5 Pro Realtime"] STT --> TD["Turn detection
Pipecat (VAD + Smart Turn)
or AssemblyAI"] TD --> LLM["LLM"] LLM --> TTS["TTS"] TTS --> U ``` Once you have an agent running, tune it for what matters most to your use case: Decide when the user is done speaking — the two Pipecat modes, defaults, and entity tuning. Shorten the gap between the user finishing and the agent replying. Prompting, key terms, conversation context, and noise handling. Natural barge-in while the agent is speaking. } href="https://docs.pipecat.ai/server/services/stt/assemblyai" > View Pipecat's AssemblyAI STT plugin reference. For a standalone voice agent without Pipecat or an external LLM, see the [AssemblyAI Voice Agent API](/voice-agents/voice-agent-api), which handles STT, LLM routing, and TTS in a single WebSocket connection. ## Quickstart Get a working, talking agent in a few minutes, then optimize from there. Install Pipecat with the AssemblyAI, LLM, and TTS extras you need: ```bash pip install "pipecat-ai[assemblyai,openai,cartesia]" python-dotenv ``` **What's included:** - `assemblyai`: AssemblyAI U3 Pro STT service - `openai`: OpenAI LLM service (used in the example) - `cartesia`: Cartesia TTS service (used in the example) The example uses OpenAI and Cartesia, but you can use any LLM or TTS supported by Pipecat — just swap the extras (e.g., `pipecat-ai[assemblyai,anthropic,elevenlabs]`). Universal 3.5 Pro Realtime, automatic [conversation context](#conversation-context), and [Voice Focus](#voice-focus) require **`pipecat-ai` 1.4.0+**. Older versions won't recognize the `universal-3-5-pro` model. Set your API keys in a `.env` file: ```env ASSEMBLYAI_API_KEY=your_assemblyai_key OPENAI_API_KEY=your_openai_key CARTESIA_API_KEY=your_cartesia_key ``` You can obtain an AssemblyAI API key by [signing up for a free account](https://www.assemblyai.com/dashboard/signup) and navigating to the [API Keys tab](https://www.assemblyai.com/dashboard/home) of the dashboard. The example below uses Pipecat-controlled turn detection (the default). Pay attention to the comments for switching to AssemblyAI's built-in turn detection, and note that the **assistant aggregator** at the end of the pipeline is what enables automatic [conversation context](#conversation-context). ```python expandable import os from dotenv import load_dotenv from loguru import logger from pipecat.audio.vad.silero import SileroVADAnalyzer from pipecat.frames.frames import LLMRunFrame from pipecat.pipeline.pipeline import Pipeline from pipecat.pipeline.worker import PipelineParams, PipelineWorker from pipecat.processors.aggregators.llm_context import LLMContext from pipecat.processors.aggregators.llm_response_universal import ( LLMContextAggregatorPair, LLMUserAggregatorParams, ) from pipecat.runner.types import RunnerArguments from pipecat.runner.utils import create_transport from pipecat.services.assemblyai.stt import AssemblyAISTTService from pipecat.services.cartesia.tts import CartesiaTTSService from pipecat.services.openai.llm import OpenAILLMService from pipecat.transports.base_transport import BaseTransport, TransportParams from pipecat.transports.daily.transport import DailyParams from pipecat.workers.runner import WorkerRunner load_dotenv() transport_params = { "daily": lambda: DailyParams(audio_in_enabled=True, audio_out_enabled=True), "webrtc": lambda: TransportParams(audio_in_enabled=True, audio_out_enabled=True), } async def run_bot(transport: BaseTransport, runner_args: RunnerArguments): stt = AssemblyAISTTService( api_key=os.environ["ASSEMBLYAI_API_KEY"], settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", min_turn_silence=100, # max_turn_silence is auto-synced to min_turn_silence in Pipecat mode. # vad_threshold=0.3, # Align with your local VAD's threshold # continuous_partials=True, # Default — steady ~3s partials during long turns # interruption_delay=0, # Optional: faster first partial (~300ms effective) ), vad_force_turn_endpoint=True, # Pipecat mode (default). # Set False to use AssemblyAI's built-in turn detection (universal-3-5-pro / universal-3-5-pro only): # vad_force_turn_endpoint=False, ) llm = OpenAILLMService(api_key=os.environ["OPENAI_API_KEY"]) tts = CartesiaTTSService(api_key=os.environ["CARTESIA_API_KEY"]) context = LLMContext() user_aggregator, assistant_aggregator = LLMContextAggregatorPair( context, user_params=LLMUserAggregatorParams(vad_analyzer=SileroVADAnalyzer()), ) pipeline = Pipeline( [ transport.input(), # Transport user input stt, # STT user_aggregator, # User responses llm, # LLM tts, # TTS transport.output(), # Transport bot output assistant_aggregator, # Assistant responses → automatic conversation context ] ) worker = PipelineWorker(pipeline, params=PipelineParams(enable_metrics=True)) @transport.event_handler("on_client_connected") async def on_client_connected(transport, client): context.add_message( {"role": "system", "content": "You are a helpful voice assistant. Keep replies brief and speakable."} ) await worker.queue_frames([LLMRunFrame()]) runner = WorkerRunner(handle_sigint=runner_args.handle_sigint) await runner.add_workers(worker) await runner.run() async def bot(runner_args: RunnerArguments): transport = await create_transport(runner_args, transport_params) await run_bot(transport, runner_args) if __name__ == "__main__": from pipecat.runner.run import main main() ``` Two complete, runnable examples live in the Pipecat repo: [voice-assemblyai.py](https://github.com/pipecat-ai/pipecat/blob/main/examples/voice/voice-assemblyai.py) (Pipecat turn detection) and [voice-assemblyai-turn-detection.py](https://github.com/pipecat-ai/pipecat/blob/main/examples/voice/voice-assemblyai-turn-detection.py) (AssemblyAI's built-in turn detection). Run the agent directly with local audio: ```bash python your_agent.py ``` Speak into your microphone after hearing the greeting. For WebRTC or Daily testing, see [Running your agent](#running-your-agent). ## Parameters reference ### **Universal 3.5 Pro Realtime** parameters These are the key parameters to tune. Set them inside `AssemblyAISTTService.Settings(...)`. They apply to the whole U3 Pro family (`universal-3-5-pro` and `universal-3-5-pro`). The streaming model. `"universal-3-5-pro"` is the recommended flagship model; the plugin currently defaults to `"universal-3-5-pro"`, so set `model` explicitly. Both belong to the U3 Pro family and share every parameter below. Accuracy/latency preset: `"min_latency"`, `"balanced"`, or `"max_accuracy"`. Sets sensible defaults for mode-dependent fields; any value you set explicitly still takes precedence. The server defaults to `"balanced"`. Construction-time only. U3 Pro family only. See [Optimizing accuracy and latency](/streaming/getting-started/optimizing-accuracy-and-latency). List of terms to boost recognition for. Used on its own, your terms are appended to the default prompt automatically. Can't be set in the same request as `prompt` — see [Key terms](#key-terms) to combine boosting with a custom prompt. Contextual prompt — a natural-language description of what the audio is about (domain, scenario, or full details). Can't be set in the same request as `keyterms_prompt`; fold the terms into the prompt text instead (see [Key terms](#key-terms)). **Prompting is currently a beta feature**: see [Prompting](/streaming/prompting-and-keyterms) for more information. Context carryover seed — your agent's most recent spoken reply, up to ~1500 characters, used to transcribe the next user turn more accurately. Set it at construction time to seed an opening greeting; later turns are fed automatically. U3 Pro family only. See [Conversation context](#conversation-context). How many prior conversation entries are carried forward automatically. Range `0`–`100`; `0` disables carryover entirely (including the automatic `agent_context` feed). Construction-time only; leave unset for the server default (`5` on `universal-3-5-pro`; `3` on older `u3-rt-pro`). U3 Pro family only. Milliseconds of silence before a speculative end-of-turn check. When the check fires, the model looks for terminal punctuation (`.` `?` `!`) to decide whether the turn has ended. (Formerly `min_end_of_turn_silence_when_confident`, deprecated but still supported with a warning.) Maximum silence before the turn is forced to end, regardless of punctuation. **Auto-synced to `min_turn_silence` in Pipecat mode**; respected as configured in AssemblyAI's built-in turn detection mode. AssemblyAI's internal VAD threshold (`0.0`–`1.0`) for classifying audio frames as silence. Align with your local VAD's activation threshold to avoid a "dead zone" where AssemblyAI transcribes speech your VAD hasn't detected yet. Server-side noise suppression that isolates the primary speaker. `"near-field"` for close-talking mics, `"far-field"` for distant capture. Construction-time only. U3 Pro family only. See [Voice focus](#voice-focus). How aggressively `voice_focus` suppresses background audio. `0.0`–`1.0`; higher is more aggressive. Only takes effect when `voice_focus` is set. Construction-time only. U3 Pro family only. Whether to emit additional partial transcripts during long turns at a steady ~3 second cadence. When enabled (default on both the API and this plugin), additional partials covering the full turn transcript are emitted approximately every 3 seconds while speech continues. When disabled, only one early partial is emitted near turn start. The first partial (at 750ms) is unaffected. Useful when downstream consumers (LLMs, UI, eager inference) need frequent updates during long, uninterrupted turns. See [Continuous partials](/streaming/getting-started/transcribe-streaming-audio) for details. How soon the first partial transcript is emitted during a turn, in milliseconds. Range: `0`–`1000`. Lower values produce faster time to first token (TTFT) for barge-in and speculative inference; higher values produce more confident first partials. The server adds a minimum of 300ms on top of the configured value (`interruption_delay=0` → ~300ms effective, `interruption_delay=500` → ~800ms effective). See [Tuning early partial timing](/streaming/getting-started/transcribe-streaming-audio) for details. **Universal 3.5 Pro Realtime** code-switches natively between supported languages. This parameter controls whether `language_code` and `language_confidence` are included in turn messages. Enable speaker diarization. See [Speaker diarization](#speaker-diarization). ### General parameters These apply across models and Pipecat setups. `api_key`, `vad_force_turn_endpoint`, `should_interrupt`, and `speaker_format` are passed directly to `AssemblyAISTTService(...)`, not inside `Settings`. Your AssemblyAI API key. `True` for Pipecat mode (VAD + Smart Turn controls turns); `False` for AssemblyAI's built-in turn detection (`universal-3-5-pro` / `universal-3-5-pro` only). See [Turn detection](#turn-detection). Whether the user starting to speak interrupts the bot. Only applies in AssemblyAI's built-in turn detection mode (`vad_force_turn_endpoint=False`). Template string for formatting speaker labels (e.g., `"[{speaker}] {text}"`). Used with `speaker_labels`. The sample rate of the audio stream. The encoding of the audio stream. Allowed values: `pcm_s16le`, `pcm_mulaw`. ### Legacy parameters These apply to the `universal-streaming-english` and `universal-streaming-multilingual` models, but **do not affect Universal 3.5 Pro Realtime or `universal-3-5-pro`**: Confidence threshold for end-of-turn detection. The U3 Pro family uses punctuation-based turn detection instead, so this parameter has no effect. Whether to return formatted final transcripts. The U3 Pro family always returns formatted transcripts, so this parameter no longer applies. ## Turn detection In Pipecat, you choose **which component decides when the user is done speaking** with the `vad_force_turn_endpoint` flag on `AssemblyAISTTService`. The U3 Pro family uses a **punctuation-based** end-of-turn system: after a period of silence, the model checks for terminal punctuation (`.` `?` `!`) rather than a confidence score. For more on how this works, see [Configuring turn detection](/streaming/getting-started/transcribe-streaming-audio). The `vad_force_turn_endpoint` parameter controls which turn detection mode is used. It defaults to `True` (Pipecat mode), which sends a `ForceEndpoint` message to AssemblyAI when the local VAD detects silence. Set it to `False` to use AssemblyAI's built-in turn detection instead. Choosing the right mode is critical for balancing responsiveness and turn accuracy in your voice agent. ### Pipecat mode (default, recommended) **When to use:** Most voice agent applications requiring responsive interruptions. ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", min_turn_silence=100, ), vad_force_turn_endpoint=True, # Default (Pipecat mode) ) ``` **How it works:** - VAD + the Smart Turn analyzer control when the user is done speaking. - A `ForceEndpoint` message is sent to AssemblyAI on VAD silence detection. - `max_turn_silence` is **automatically synchronized** with `min_turn_silence`. - Best for low-latency, responsive voice agents. ### AssemblyAI's built-in turn detection **When to use:** When you want AssemblyAI's punctuation-based turn detection to control turn endings, configured through the settings below. ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", min_turn_silence=100, max_turn_silence=1000, # Now respected independently ), vad_force_turn_endpoint=False, # AssemblyAI's built-in turn detection ) ``` **How it works:** 1. User speaks → audio streams to AssemblyAI. 2. User pauses for `min_turn_silence` (e.g., `100ms`) → the model checks for terminal punctuation. 3. If terminal punctuation (`.` `?` `!`) is found → the turn ends immediately. 4. If not → a partial is emitted and the turn continues waiting. 5. If silence reaches `max_turn_silence` (e.g., `1000ms`) → the turn is forced to end regardless. In this mode all timing parameters are respected as configured, the service emits `UserStartedSpeakingFrame` / `UserStoppedSpeakingFrame`, and `SpeechStarted` events drive fast barge-in. Only available with `universal-3-5-pro` / `universal-3-5-pro` (other models require Pipecat mode). ### Entity splitting tradeoff Lower `min_turn_silence` and `max_turn_silence` values produce faster transcripts but can split entities or utterances across turns. The two parameters affect this differently. #### `min_turn_silence` too low The speculative check fires too early, splitting entities on punctuation: ```text # With (min_turn_silence=100, max_turn_silence=1000) "It's John." → FINAL (100ms pause, check fires, period found → turn ends) "Smith." → FINAL "At gmail.com." → FINAL # With (min_turn_silence=400, max_turn_silence=1000) "It's john.smith@gmail.com." → FINAL (single turn, properly formatted) ``` #### `max_turn_silence` too low The forced turn-end cuts off the user mid-thought: ```text # With (min_turn_silence=100, max_turn_silence=1000) "I wanted to check on my order from..." → FINAL (1000ms silence, forced end) "last Tuesday, order number 4829." → FINAL (new turn) # With (min_turn_silence=100, max_turn_silence=2000) "I wanted to check on my order from last Tuesday, order number 4829." → FINAL (single turn) ``` **Universal 3.5 Pro Realtime's** formatting is significantly better when it has full context in a single turn — email addresses, phone numbers, credit card numbers, and physical addresses all benefit. If your use case involves alphanumeric dictation, raise `max_turn_silence` during those portions of the conversation (e.g., to `2000`–`4000` ms) using [dynamic configuration](#dynamic-configuration), then lower it again afterward. In Pipecat mode, raise `min_turn_silence` (which `max_turn_silence` follows) for the same effect. ## Latency A voice agent feels responsive when the gap between the user finishing and the agent replying is short. Start with the **`mode` preset** — the highest-level dial for the accuracy/latency trade-off. It sets sensible defaults for the fine-grained levers below, so you can pick a target and tune from there: ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", mode="balanced", # "min_latency" (fastest) · "balanced" · "max_accuracy" (best quality) ), ) ``` `mode` is set at construction time (it can't be changed mid-session) and influences the defaults of the levers below. Any value you set explicitly still wins. Leave it unset to use the server's default preset. See [Optimizing accuracy and latency](/streaming/getting-started/optimizing-accuracy-and-latency). From there, fine-tune the individual levers: - **End-of-turn timing.** `min_turn_silence` (speculative check) and `max_turn_silence` (forced end) directly control how soon a turn ends. Lower is faster but risks splitting entities — see [Turn detection](#turn-detection). - **Time to first partial.** `interruption_delay` controls how soon the first partial is emitted, which drives faster barge-in and speculative inference. The server adds a minimum of `300ms` on top of the configured value. - **Sample rate.** Use 16 kHz (`sample_rate=16000`). Higher rates don't improve accuracy and only add bandwidth. - **Continuous partials.** `continuous_partials` (on by default) emits a partial every ~3 seconds during long turns. Leave it on for steady mid-turn updates, or disable it if you only need a single early partial. - **Skip client-side preprocessing.** Don't run your own noise cancellation before audio reaches the model — the artifacts it introduces usually hurt accuracy more than the original noise. Use server-side [Voice Focus](#voice-focus) instead. ### Latency breakdown | Stage | Typical | Controlled by | | ---------------------------------------- | ---------------------------------------------------- | -------------------- | | Network round trip | ~50 ms | — | | Speech-to-text | ~200–300 ms | model | | First partial (TTFT) | configured `interruption_delay` + ~300 ms server min | `interruption_delay` | | End of turn (terminal punctuation found) | `min_turn_silence` (default 100 ms) | `min_turn_silence` | | End of turn (no punctuation, forced) | up to `max_turn_silence` | `max_turn_silence` | ## Accuracy **Universal 3.5 Pro Realtime** is accurate out of the box. When you need more — domain vocabulary, proper nouns, noisy audio — reach for these levers. For entity-heavy dictation, also tune turn detection (see [Entity splitting tradeoff](#entity-splitting-tradeoff)), and note that the high-level [`mode` preset](#latency) shifts the overall accuracy/latency balance (use `max_accuracy` to favor quality). ### Prompting **Beta feature** Prompting is considered a beta feature for **Universal 3.5 Pro Realtime**. While it can be a powerful tool for improving accuracy in certain use cases, **we recommend starting without a `prompt` to first establish baseline performance.** Once the baseline has been tested, you can add context to further optimize for your use case (e.g., language mix to expect, use case or domain). **Universal 3.5 Pro Realtime** supports a `prompt` parameter for [contextual prompting](/streaming/prompting-and-keyterms) — a description of what the audio is about. Transcription behavior (verbatim output, punctuation, turn detection) is built in and optimized automatically; the prompt carries context, not instructions. ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", prompt="Customer support call about an internet service outage.", ), ) ``` ### Key terms Use `keyterms_prompt` to boost recognition of specific names, brands, or domain terms. On its own, your terms are appended to the default prompt automatically — so you get boosting and prompting together: ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", keyterms_prompt=["Xiomara", "Saoirse", "Pipecat", "AssemblyAI"], ), ) ``` You can't pass `prompt` and `keyterms_prompt` in the same request — doing so raises a validation error. You don't have to give up term boosting to use a contextual prompt, though. Either: - Pass **`keyterms_prompt` on its own** — your terms are appended to the default prompt automatically, or - Fold the terms into a custom **`prompt`**, e.g. end it with `"Make sure to boost the words Xiomara, Saoirse, Pipecat in the audio."` ### Conversation context Give the model both sides of the dialog so it transcribes the next user turn more accurately. **Universal 3.5 Pro Realtime** keeps a short, per-session memory of the conversation from two sources: - **The agent half** — what your agent just said. - **The user half** — prior STT-finalized user turns. With the agent's question in context, the model can anticipate the answer, sharpen entity recognition, and disambiguate similar-sounding words. For example, after your agent asks `"What's your email address?"`, the model can produce `"user@assemblyai.com"` instead of `"user at assemblyai dot com"`. This has the biggest impact on short replies (`"yes"`, `"7pm"`, single names) and spelled-out entities. See [Conversation context](/streaming/universal-3-5-pro/context-carryover) for the full reference. **In Pipecat, conversation context is automatic — no event wiring required.** As long as your pipeline includes the standard LLM context aggregator (the `assistant_aggregator` from `LLMContextAggregatorPair`), Pipecat broadcasts an `LLMContextAssistantTurnFrame` when each bot turn completes, and `AssemblyAISTTService` feeds that reply to the model as `agent_context` automatically. Just use a U3 Pro family model on `pipecat-ai` 1.4.0+. | Parameter | Type | Description | | -------------------------- | ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `agent_context` | str | Your agent's most recent spoken reply, up to ~1500 characters. Set it at construction time to seed an opening greeting; subsequent replies are fed automatically. | | `previous_context_n_turns` | int | How many prior conversation entries are carried forward automatically. Range `0`–`100`; `0` disables carryover entirely. Construction-time only; server default is `5` on `universal-3-5-pro` (`3` on older `u3-rt-pro`). | #### Seeding the opening greeting The automatic feed kicks in once your agent completes its first turn. To give the model context for the user's *very first* reply (the answer to your greeting), set `agent_context` at construction time: ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", # Seed the opening line; later turns are fed automatically by the aggregator. agent_context="Hi! Thanks for calling Acme. What's the email on your account?", # previous_context_n_turns=5, # Default on universal-3-5-pro. Set 0 to disable carryover entirely. ), ) ``` #### Manual control with `update_agent_context()` If your pipeline doesn't use the standard LLM context aggregator, or you want explicit control over what the model sees, push the agent's reply yourself. This is a **live update — no reconnect required**: ```python # Whenever your agent finishes speaking: await stt.update_agent_context("Your account is past due. Would you like to pay now?") ``` `agent_context`, `previous_context_n_turns`, and `update_agent_context()` are supported only on the U3 Pro family (`universal-3-5-pro`, `universal-3-5-pro`). Values are clipped to ~1500 characters and re-seeded automatically on reconnect. Setting `previous_context_n_turns=0` disables the automatic feed as well. ### Voice focus Voice Focus isolates the primary speaker and suppresses background noise — chatter, keyboard clicks, fan hum, room echo — **server-side, before audio reaches the model**. Use it instead of client-side noise cancellation, which tends to introduce artifacts that hurt accuracy more than the noise itself. | Parameter | Type | Description | | ----------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------ | | `voice_focus` | str | `"near-field"` for headsets, handsets, and other close-talking mics; `"far-field"` for conference rooms, laptop mics, and other distant capture. | | `voice_focus_threshold` | float | Optional. `0.0`–`1.0`; higher values suppress background audio more aggressively. | Both are construction-time parameters on the U3 Pro family. See [Voice Focus](/streaming/voice-focus) for details. ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", voice_focus="far-field", # "near-field" for close-talking mics voice_focus_threshold=0.5, # Optional: 0.0–1.0, higher = more aggressive ), ) ``` ## Interruption handling Barge-in — the user interrupting while the agent is speaking — is handled by Pipecat, and the signals that drive it depend on your turn detection mode. - **Pipecat mode (`vad_force_turn_endpoint=True`).** Pipecat's local VAD and the Smart Turn analyzer detect the user starting to speak and interrupt the bot's TTS. AssemblyAI also emits `SpeechStarted` events as a backstop. - **AssemblyAI's built-in turn detection (`vad_force_turn_endpoint=False`).** The service emits `UserStartedSpeakingFrame` / `UserStoppedSpeakingFrame` and uses AssemblyAI's `SpeechStarted` events for fast barge-in. Set `should_interrupt=False` (constructor argument) to disable barge-in entirely in this mode. ```json {"type": "SpeechStarted", "timestamp": 14400, "confidence": 0.79} ``` On detection, Pipecat stops TTS playback and switches to listening. To reduce false interruptions from short backchannels (`"mhm"`, `"yeah"`, `"okay"`), keep your VAD threshold aligned with `vad_threshold` and lean on Pipecat's Smart Turn analyzer, which evaluates whether speech is a genuine turn rather than a filler. ## Dynamic configuration Update settings mid-conversation by queueing an `STTUpdateSettingsFrame` with a settings delta — adapt to the conversation stage as it unfolds. See [stt-assemblyai.py](https://github.com/pipecat-ai/pipecat/blob/main/examples/update-settings/stt/stt-assemblyai.py) for a complete working example. ```python expandable from pipecat.frames.frames import STTUpdateSettingsFrame from pipecat.services.assemblyai.stt import AssemblyAISTTService # Update keyterms during the conversation await worker.queue_frame( STTUpdateSettingsFrame( delta=AssemblyAISTTService.Settings( keyterms_prompt=["NewName", "NewCompany"], ) ) ) # Widen the silence window during entity dictation await worker.queue_frame( STTUpdateSettingsFrame( delta=AssemblyAISTTService.Settings( min_turn_silence=200, max_turn_silence=3000, # Respected in AssemblyAI's built-in turn detection mode ) ) ) ``` **`agent_context` is the only setting applied live.** Changing any other setting via `STTUpdateSettingsFrame` reconnects the AssemblyAI session to apply it (a brief interruption). To push conversation context without a reconnect, use the dedicated `stt.update_agent_context(...)` method — see [Conversation context](#conversation-context). | Conversation stage | Adjustment | | ------------------------------------------ | -------------------------------------------------------------------------------- | | Caller identification (names, account IDs) | Boost terms with `keyterms_prompt` | | Entity dictation (email, phone, address) | Raise `max_turn_silence` to ~`2000`–`4000` ms, then lower it again afterward | | After each agent reply | Automatic — or push `agent_context` via `update_agent_context()` | | Faster barge-in | Lower `interruption_delay` | For more information, see [Updating configuration mid-stream](/streaming/updating-configuration-mid-stream). ## Speaker diarization Identify different speakers in multi-party conversations. ### Basic diarization ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", speaker_labels=True, ), ) ``` Speaker labels (e.g., `"A"`, `"B"`, `"C"`) are included in final transcripts. ### With custom formatting Format transcripts with speaker labels for LLM context: ```python stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), settings=AssemblyAISTTService.Settings( model="universal-3-5-pro", speaker_labels=True, ), speaker_format="<{speaker}>{text}", ) ``` **Format options:** | Style | Format string | | -------- | ------------------------------- | | XML | `<{speaker}>{text}` | | Markdown | `**{speaker}**: {text}` | | Bracket | `[{speaker}] {text}` | ## Running your agent ### Development mode (local audio) ```bash python your_agent.py ``` Speak into your microphone after hearing the greeting. ### Production with Daily For production deployments, use the [Daily transport](https://docs.pipecat.ai/server/services/transport/daily) for WebRTC-based real-time audio/video. Your agent joins a Daily room as a participant and handles audio I/O through Daily's infrastructure. ### Telephony with Telnyx When bridging phone calls through Pipecat (e.g., via Telnyx), the audio is 8 kHz, not 16 kHz. Match the transport sample rates: ```python transport = TelnyxTransport( # ... audio_in_sample_rate=8000, audio_out_sample_rate=8000, ) ``` ## Troubleshooting | Issue | Cause | Solution | | ------------------------------------ | ---------------------------------------------------- | ------------------------------------------------------------------------------------------------- | | `universal-3-5-pro` not recognized | `pipecat-ai` older than 1.4.0 | Upgrade: `pip install -U "pipecat-ai[assemblyai]"` | | Turn over-segmentation | `min_turn_silence` too low | Increase from `100` to `200`–`500` | | Entities split across turns | `max_turn_silence` too low (AssemblyAI mode) | Increase `max_turn_silence` (e.g., `1500`–`3500`); in Pipecat mode, raise `min_turn_silence` | | Latency on non-terminal utterances | `max_turn_silence` too high | Lower `max_turn_silence` | | Conversation context has no effect | Non-U3-Pro model, or `previous_context_n_turns=0` | Use a U3 Pro family model and leave `previous_context_n_turns` unset (or > 0) | | Mid-session setting change drops audio | Reconnect on a non-`agent_context` setting change | Expected — only `agent_context` updates live; use `update_agent_context()` for context | | Mis-heard names, brands, or jargon | No vocabulary hints | Add `keyterms_prompt`, or supply `prompt`/`agent_context` for context | | Poor accuracy in noisy audio | Background noise or room echo | Enable `voice_focus` (`near-field` or `far-field`) | ## Migrating from another STT provider To balance accuracy, latency, turn-taking, and interruption handling, map your current setup to AssemblyAI using the questions below. ### How are you detecting end-of-turn today? | Today | Recommended on AssemblyAI | | --------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | Your STT provider's own end-of-turn model | AssemblyAI's built-in turn detection: `vad_force_turn_endpoint=False` with `min_turn_silence=100`, `max_turn_silence=1000`. | | Silence / VAD only, with your own turn logic | Pipecat mode (`vad_force_turn_endpoint=True`, default). VAD + Smart Turn decide turns; AssemblyAI returns finals ASAP. | | You want the framework to own turn-taking | Pipecat mode (default) — Pipecat's Smart Turn analyzer makes the turn decision. | ### Which model and settings are you migrating from? | What you pass today | AssemblyAI equivalent | | ------------------------------------------ | ------------------------------------------------------------------------------------------------------ | | Current model (Deepgram, ElevenLabs, etc.) | `model="universal-3-5-pro"` (recommended flagship) or `"universal-3-5-pro"` | | Overall accuracy/latency tuning | `mode="min_latency"` / `"balanced"` / `"max_accuracy"` — a one-line starting point before fine-tuning | | Endpointing / silence thresholds | `min_turn_silence` (speculative end-of-turn) and `max_turn_silence` (forced end) | | Custom vocabulary / keywords | `keyterms_prompt=[...]`; broader domain context → `prompt` | | Provider-side conversation context | Automatic — include the LLM context aggregator; seed greetings via `agent_context` | | Formatting / punctuation toggles | On by default — formatted transcripts always (`format_turns` does not apply) | | Telephony / SIP routing | `sample_rate=8000` and `encoding="pcm_mulaw"` for 8 kHz telephony | | Client-side noise cancellation | Drop it; use server-side [Voice Focus](#voice-focus) instead | Migrating a production deployment? [Talk to our team](https://www.assemblyai.com/contact/sales). ## Speech model comparison Interested in using a different model? | Feature | U3 Pro family
(`universal-3-5-pro`, `universal-3-5-pro`) | universal-streaming-english | universal-streaming-multilingual | | ---------------------------------- | ----------------------------------------------------- | --------------------------- | -------------------------------- | | **_Turn Detection Modes_** | | | | | Pipecat mode (VAD + Smart Turn) | ✅ | ✅ | ✅ | | AssemblyAI turn detection mode | ✅ | ❌ | ❌ | | **_Turn Detection Parameters_** | | | | | `min_turn_silence` | ✅ | ✅ | ✅ | | `max_turn_silence` | ✅ | ✅ | ✅ | | `end_of_turn_confidence_threshold` | ❌ | ✅ (1.0) | ✅ (1.0) | | `continuous_partials` | ✅ | ❌ | ❌ | | `interruption_delay` | ✅ | ❌ | ❌ | | **_Advanced Features_** | | | | | Keyterms boosting | ✅ | ✅ | ✅ | | Custom prompting (beta) | ✅ | ❌ | ❌ | | Conversation context (carryover) | ✅ | ❌ | ❌ | | Voice Focus | ✅ | ❌ | ❌ | | Speaker diarization | ✅ | ✅ | ✅ | | Dynamic parameter updates | ✅ | ✅ | ✅ | | **_Language Support_** | | | | | Multilingual code switching | ✅ | ❌ | ✅ | | Language detection | ✅ | ❌ | ✅ | **Legend:** - ✅ Fully supported and recommended - ❌ Not supported / Not used **The U3 Pro family is recommended** for all new voice agent implementations. The universal-streaming models are maintained for backward compatibility but lack the optimizations and features specifically designed for real-time conversational AI. The `end_of_turn_confidence_threshold` parameter is **not used** with the U3 Pro family (it won't affect behavior). For universal-streaming models, Pipecat automatically sets it to `1.0` in Pipecat mode to disable semantic turn detection and ensure fast responses. You don't need to configure this parameter manually. --- # Zapier Integration with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/zapier Source: docs/integrations/zapier.mdx Navigation: Overview > Use cases & integrations > Integrations > No-code tools Description: Transcribe audio in Zapier using the AssemblyAI app. Zapier is a workflow automation tool that lets you integrate various services together without requiring coding knowledge. You can use our AI models to process audio data by transcribing it with speech recognition models and analyzing it with speech understanding models. You can supply audio to the AssemblyAI app and connect the output of our models to other services in your Zaps. ## Quickstart In your Zap editor, add an action, search for `AssemblyAI` and select the AssemblyAI app. ![Change action Zapier screen, with AssemblyAI in search box.](/assets/img/integrations/zapier/1-change-action.png) Next, configure the action. In the **App & event** tab, select **Transcribe** for the **Event** dropdown, then click **Continue**. ![App & event Zapier tab with Event field set to "Transcribe".](/assets/img/integrations/zapier/2-app-event.png) Then, in the **Account** tab, click **Sign in** which will open a separate window. In the window, enter your AssemblyAI API key in **API Key** field, and click **Yes, Continue to AssemblyAI**. Back in the Zap editor, click **Continue**. ![Account Zapier tab where you are prompted to Connect AssemblyAI with a Sign in button.](/assets/img/integrations/zapier/3-account.png) In the **Action** tab, enter the URL of the audio or video file you want to transcribe in the **Audio URL** field. The URL has to be publicly accessible. Click **Continue**. ![Action Zapier tab where you are prompted to enter the Audio URL to transcribe using AssemblyAI.](/assets/img/integrations/zapier/4-action.png) Finally, you can test the action. You can use all the fields returned by the action in subsequent steps. All AssemblyAI actions return sample data during testing instead of running the action. This makes it easier to build your Zaps, however, you have to test using normal Zap runs to verify everything is working correctly. Learn more about why we [return sample data during testing below](#testing-with-sample-data). ![Test Zapier tab where you can see the output of the AssemblyAI Transcribe action.](/assets/img/integrations/zapier/5-test.png) ## Zapier Actions ### Transcribe Transcribe an audio file and wait until the transcript has completed or failed. Configure the `Audio URL` field with the URL of the audio file you want to transcribe. The `Audio URL` must be accessible by AssemblyAI's servers. If you don't want to wait until the transcript is ready, change the `Wait until Transcript is Ready` parameter to `False`. ### Get Transcript Retrieves a transcript by its ID. ### Get Subtitles Export the transcript as SRT or VTT subtitles. You can only invoke this action after the transcript is completed. ### Get Sentences Retrieve the sentences of the transcript by its ID. You can only invoke this action after the transcript is completed. ### Get Paragraphs Retrieve the paragraphs of the transcript by its ID. You can only invoke this action after the transcript is completed. ## Testing with sample data A transcript goes through multiple phases to transcribe audio, reflected by different statuses. The initial status is typically `processing`, and the final status is either `completed` or `error`. A transcript may also have a `queued` status if the job is waiting to be processed (for example, when you've exceeded your rate limit). **During a normal Zap run**, the Transcribe event will wait until the transcript status is `completed`, and throw an error if the status is `error`. Unfortunately, this is not the case during testing. Because of a Zapier platform limitation, **during testing**, the Transcribe event will return a transcript before it has reached the `completed` status. A transcript that does not have the `completed` status cannot be used in subsequent tests. This way you can easily test using sample data, but you still have to use normal Zap runs to verify the end-to-end functionality. ## Additional resources You can learn more about using Zapier with AssemblyAI in these resources: - [How to generate subtitles for your videos using the AssemblyAI app for Zapier](https://www.assemblyai.com/blog/generate-subtitles-with-zapier) - [How to get started with AssemblyAI on Zapier](https://help.zapier.com/hc/en-us/articles/16411509681933-How-to-get-started-with-AssemblyAI-on-Zapier) - [AssemblyAI app on Zapier](https://zapier.com/apps/assemblyai/integrations) --- # Integrate Make with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/make Source: docs/integrations/make.mdx Navigation: Overview > Use cases & integrations > Integrations > No-code tools Description: Use our Make (formerly Integromat) app to use AssemblyAI's speech AI in your Make scenarios. [Make](https://make.com/) (formerly Integromat) is a workflow automation tool that lets you integrate various services together without requiring coding knowledge. With the AssemblyAI app for Make, you can use our AI models to process audio data by transcribing it with speech recognition models, analyzing it with speech understanding models, and building generative features on top of it with LLMs. You can supply audio to the AssemblyAI app and connect the output of our models to other services in your Make scenarios. ## Quickstart Create or edit a scenario in Make. Add a new module, search for AssemblyAI, and select the module that you want to use. ![Search for AssemblyAI modules in Make](/assets/img/integrations/make/search-module.png) Select the module that you want to use. ![Add an AssemblyAI module in Make](/assets/img/integrations/make/add-module.png) Create a new connection or select an existing one. In **AssemblyAI API Key**, enter the API key from your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home). In **Environment**, select the [region you would like to use](/pre-recorded-audio/select-the-region). You must select either US or EU. Click **Save**. ![Create a connection to AssemblyAI in Make](/assets/img/integrations/make/create-connection.png) Finally, configure your AssemblyAI module. Continue reading to learn more about all the available modules. ![Configure an AssemblyAI module in Make](/assets/img/integrations/make/configure-module.png) ## AssemblyAI app modules The AssemblyAI app for Make provides the following modules: ### Files #### Upload a File Upload an audio file to AssemblyAI so you can transcribe it. You can pass the `Upload URL` output field to the `Audio URL` input field of [Transcribe an Audio File](#transcribe-an-audio-file) module. ### Transcripts #### Transcribe an Audio File Transcribe an audio file and wait until the transcript has completed or failed. Configure the `Audio URL` field with the URL of the audio file you want to transcribe. The `Audio URL` must be accessible by AssemblyAI's servers. If you don't have a publicly accessible URL, you can use the [Upload a File](#upload-a-file) module to upload the audio file to AssemblyAI. If you don't want to wait until the transcript is ready, change the `Wait until Transcript is Ready` parameter to `No` under **Show advanced settings**. Configure your desired [Speech Understanding models](/speech-understanding) when you create the transcript. The results of the models will be included in the transcript output. #### Wait until Transcript is Ready Wait for an existing transcript to be ready. This module will complete when the status of the transcript changes to "completed" or "error". #### Watch for Transcript Ready Notification Create a webhook URL to receive a notification when a transcript is ready. When the transcript is ready, the webhook will be invoked with the transcript status and ID. The status will be "completed" or "error". #### Get a Transcript Retrieve a transcript by ID. #### Get Paragraphs of a Transcript Retrieve the paragraphs of a transcript. You can only invoke this module after the transcript is completed. #### Get Sentences of a Transcript Retrieve the sentences of a transcript. You can only invoke this module after the transcript is completed. #### Get Subtitles for a Transcript Create SRT or VTT subtitles for a transcript. You can only invoke this module after the transcript is completed. #### Get Redacted Audio of a Transcript First, you need to configure PII audio redaction using these fields when you create the transcript: - `Redact PII`: `Yes` - `Redact PII Audio`: `Yes` - `Redact PII Policies`: Configure at least one PII policy Then, you can use this module to retrieve the redacted audio of the transcript. You can only invoke this module after the transcript is completed. #### Search for Words in a Transcript Search for words in a transcript. You can only invoke this module after the transcript is completed. #### List Transcripts Paginate over all transcripts. #### Delete a Transcript Delete a transcript by ID. Deleting a transcript does not delete the transcript resource itself, but removes the data from the resource and marks it as deleted. You can only invoke this module after the transcript status is "completed" or "error". ### Other modules #### Apply LLM Gateway to Transcripts Send a message to [LLM Gateway](/llm-gateway/chat-completions). Configure the `Model` field with the model you want to use, and the `Role` and `Content` fields to build your message. See [Available models](/llm-gateway/available-models) for the list of supported models, and [Cloud Endpoints and Data Residency](/llm-gateway/cloud-endpoints-and-data-residency) for which models are available in each region. To apply LLM Gateway to a transcript, map the `Text` output field from the [Transcribe an Audio File](#transcribe-an-audio-file) module into the `Content` field. ![Apply LLM Gateway to Transcripts module in Make](/assets/img/integrations/make/llmgw-module.png) #### Make an API Call Make your own REST API HTTP requests to the AssemblyAI API using your existing connection. ## Additional resources You can learn more about using Make with AssemblyAI in these resources: - [Iterate Over Speaker Labels in Make](/pre-recorded-audio/guides/make-speaker-labels) - [Redact PII in Audio with Make and AssemblyAI](https://www.assemblyai.com/blog/redact-pii-audio-with-make/) - [Make Integration page for AssemblyAI](https://www.make.com/en/integrations/assembly-ai) - [AssemblyAI Make App Invitation Link](https://us1.make.com/app/invite/f21437a4d43b63efc8ee9aec385d9f10) --- # The AssemblyAI n8n Integration URL: https://www.assemblyai.com/docs/integrations/n-8-n Source: docs/integrations/n-8-n.mdx Navigation: Overview > Use cases & integrations > Integrations > No-code tools Description: Integrate AssemblyAI with 1000+ apps and services using n8n's automation platform. Unlock the full potential of AssemblyAI and [n8n's](https://n8n.io/) automation platform by connecting AssemblyAI's speech-to-text and speech understanding capabilities with over 1,000 apps, data sources, services, and n8n's built-in AI features. The [AssemblyAI n8n integration](https://n8n.io/integrations/assemblyai/) is built and maintained by AssemblyAI (verified by n8n). ## Overview This comprehensive tutorial walks you through building a complete AssemblyAI workflow within n8n Cloud that: 1. Watches for changes to a Google Drive folder 2. Once an audio file is added, submits a transcription request to AssemblyAI with the audio 3. Polls for the transcription process to complete 4. Once complete, processes the transcript into a text file with just the transcript text and speaker labels 5. Saves the processed transcript `.txt` file output to a Google Drive folder 6. Deletes the transcript from AssemblyAI's servers The [Google Drive n8n integration](https://n8n.io/integrations/google-drive/) is used in this example for file storage and triggering, but you can adapt the workflow to use other services like Dropbox, Amazon S3, FTP/SFTP, or webhooks based on your needs. For all available integrations, see the [n8n integrations page](https://n8n.io/integrations/). ## Prerequisites Before you begin, you'll need: - An [AssemblyAI API key](https://www.assemblyai.com/dashboard/home) - Signing up for an AssemblyAI account is free! An account on the free tier can transcribe up to **\$50 in transcription** and will be able to process up to **five files in parallel**. If you upgrade your account before using \$50 in transcription any unused amount will be retained on your account. - An n8n account, either via [n8n Cloud](https://app.n8n.cloud/register) (recommended) or a self-hosted instance. - A [Google Cloud account](https://console.cloud.google.com/) for Google Drive API access - A folder in Google Drive to monitor for new audio files (`/audio_files`) - A folder in Google Drive to save the formatted `.txt` transcript output (`/transcripts`) ## Credentials n8n credentials page for configuring Google Drive and AssemblyAI credentials Before building your workflow, you'll need to configure credentials for both Google Drive and AssemblyAI in n8n. ### Google Drive OAuth Credentials 1. Go to your n8n credentials page 2. Click **Add Credential** and search for **"Google Drive OAuth2 API"** 3. Follow the setup process to connect your Google Drive account with Google Cloud Console: - [Video tutorial: Setting up Google Drive credentials](https://www.youtube.com/watch?v=FBGtpWMTppw&t=244s) - [n8n documentation: Google OAuth setup](https://docs.n8n.io/integrations/builtin/credentials/google/oauth-single-service/) 4. Save the credential ### AssemblyAI Credentials 1. Go to your n8n credentials page 2. Click **Add Credential** and search for **"AssemblyAI API"** 3. Enter your [AssemblyAI API key](https://www.assemblyai.com/dashboard/home) 4. Save the credential Once both credentials are configured, you can use them throughout your workflow. ## Instructions ### Step 1: Set Up Your Workflow and Trigger 1. Log in to your n8n account (either [n8n Cloud](https://app.n8n.cloud/login) or your self-hosted instance) 2. Click the **+** icon or **New Workflow** to create a new workflow 3. Give your workflow a descriptive name (e.g., "AssemblyAI Transcription Pipeline") 4. Install the AssemblyAI node: - Click the **+** icon to open the nodes panel - Search for "AssemblyAI" - Click the **Install** button next to the AssemblyAI node n8n Google Drive trigger node configuration for detecting new audio files The first step is to decide how n8n will know when your audio file is ready to transcribe. This is your workflow trigger. For this example, we'll use **Google Drive** to automatically trigger the workflow when a new audio file is added to a folder. 1. **Link your Google Drive account** with the credentials you created in the [Credentials](#credentials) step. 2. Add a **Google Drive** trigger node to your workflow 3. Select the trigger event: `On changes involving a specific folder` 4. Configure the trigger parameters: - **Credential to connect with:** Select your Google Drive credentials - **Poll Times:** Set Mode to Every Minute - **Trigger On:** Changes Involving a Specific Folder - **Folder:** From list + Select the folder in Google Drive to monitor - **Watch For:** File Created When a new audio file is added to the specified Google Drive folder, the workflow will run. **Running the workflow** To test the workflow, you can manually trigger it at any time by clicking the orange **Execute Workflow** button in the n8n editor to populate the nodes with sample data. For the Google Drive trigger to work as intended, you will need to deploy this workflow (to n8n cloud) and then add a new audio file to the monitored folder in Google Drive (`/audio_files`). This is handled in a [later step](/integrations/n-8-n#step-8-deploy). For more information on supported audio file types, size limits, and using Google Drive with AssemblyAI, see: - [Can I submit files stored in Google Drive?](/faq/can-i-submit-files-to-the-api-that-are-stored-in-a-google-drive) - [What Is the Recommended File Type for Using Your API?](/faq/what-is-the-recommended-file-type-for-using-your-api) - [What types of audio URLs can I use with the API?](/faq/what-types-of-audio-urls-can-i-use-with-the-api) - [Are there any limits on file size or duration?](/faq/are-there-any-limits-on-file-size-or-file-duration-for-files-submitted-to-the-api) **Important Google Drive Considerations** For the sake of simplicity, this tutorial uses a public Google Drive link, so you must ensure your file is **100 MB or less** and the Google Drive **folder settings are set to Public** (anyone with the link can view). Google Drive folder sharing settings set to public access If you'd like to transcribe private folders/files greater than 100MB, you'll need to add a Google Drive "Download file" node to download the binary data first, then upload it to AssemblyAI using the AssemblyAI "Upload a file" node. ### Step 2: Upload Audio File (Optional) After the Google Drive trigger fires, you have the file's `webContentLink` available at `{{ $json.webContentLink }}`. #### For files under 100MB You can **skip this step** and use the public Google Drive link directly in [Step 3](/integrations/n-8-n#step-3-submit-transcription-request) without uploading. AssemblyAI accepts [all Google Drive files below 100MB](/faq/can-i-submit-files-to-the-api-that-are-stored-in-a-google-drive). #### For files over 100MB Use the **AssemblyAI Upload** node: 1. Add a **Google Drive** "Download File" node and configure it to download the file by ID (you can get the file ID from the trigger response) 2. Add an **AssemblyAI** "Upload a file" node 3. Configure the credential with your [AssemblyAI API key](#prerequisites) 4. The upload will return an `upload_url` to use in the next step ### Step 3: Submit Transcription Request n8n AssemblyAI Create a transcription node added to the workflow Now submit the transcription request with your desired features enabled. Using the + sign linked to the Google Drive trigger, add an **AssemblyAI** `Create a transcription` node. - **Credential to connect with**: Select your AssemblyAI Account credential (if you haven't set this up yet, see [Credentials](#credentials)) - **Resource**: Transcript - **Operation**: Create - **Audio URL**: `{{ $json.webContentLink }}` - You can also drag and drop the `webContentLink` field from the Google Drive Trigger response into this field In **Additional Fields**, enable the following features for this tutorial: - **Speaker Labels**: `true` (identifies different speakers through diarization) - **Language Detection**: `true` (automatically detects the language, also known as ALD) Create a transcription node parameters with Speaker Labels and Language Detection enabled You can explore all available features, such as sentiment analysis, entity detection, pii redaction, and more in our [API Reference](/api-reference/transcripts/submit). The response will contain a `transcript_id` that you'll use to poll for completion. ### Step 4: Poll for Transcription Completion n8n workflow polling loop that checks transcription status until completion Since transcription is asynchronous, the transcription request returns immediately. You need to poll the API until the transcript is ready. Add a `Wait` node and configure the wait parameters: - **Resume**: After Time Interval - **Wait Amount**: 3.00 - **Wait Unit**: Seconds This will wait 3 seconds before checking the transcript's status. Add an **AssemblyAI** `Get a transcription` node after the `Wait` node and configure the node parameters: - **Credential to connect with**: Select your AssemblyAI Account credential (see [Credentials](#credentials)) - **Resource**: Transcript - **Operation**: Get - **Transcript ID**: `{{ $json.id }}` - You can drag and drop the `id` field from the `Create a transcription` response into this field AssemblyAI Get a transcription node parameters with Transcript ID field Add a `Switch` node after the `Get a transcription` node and configure the switch to check the transcript's [status](/api-reference/transcripts/submit#response.body.status): Add 4 routing rules, each with: - **Value 1**: `{{ $json.status }}` - Or drag the `status` field from the `Get a transcription` response - **Condition**: "is equal to" - **Value 2**: Set one of these for each rule: 1. `queued` 2. `processing` 3. `error` 4. `completed` Most transcripts go directly to `processing`. The `queued` status is only returned when a job is actually waiting to be processed, such as when you've exceeded your rate limit. The routing rule for `queued` is included so the workflow handles that case correctly when it occurs. n8n Switch node with routing rules for queued, processing, error, and completed statuses Handle each switch node output based on the transcript status: - **Output 0 (queued)**: Connect back to the **Wait** node to continue polling - **Output 1 (processing)**: Connect back to the **Wait** node to continue polling - **Output 2 (error)**: Add a **Stop and Error** node to halt the workflow - Note: You may want to add logging or resubmit the transcript instead - **Output 3 (completed)**: Continue to the next step to process the completed transcript By connecting the `processing` and `queued` outputs back to the `Get a transcription`, you create a polling loop that continues checking (every 3 seconds) until the transcript is complete or encounters an error. When the transcript is complete, the response will contain the [full transcript data](/api-reference/transcripts/submit#response). For this tutorial, we are only using the `Create a transcription`, `Get a transcription`, and `Delete a transcription` actions from the AssemblyAI node. However, n8n's AssemblyAI integration supports **all available endpoints**, including uploading files, retrieving redacted audio, getting sentences and paragraphs, LLM Gateway, and more. ### Step 5: Process Transcript Output into a Human-Readable Transcript n8n Python Code node formatting the transcript with speaker labels Once you have the completed transcript, we want to manipulate the data into a more usable format. To pull out the transcript text with speaker labels: Add a `Code` node (Python) to your workflow. Use this code to format utterances into a readable transcript: ```python transcript = _('Get a transcription').first().json formatted_text = '' if 'utterances' in transcript and transcript['utterances']: # Format with speaker labels for utterance in transcript['utterances']: formatted_text += f"Speaker {utterance['speaker']}: {utterance['text']}\n\n" else: # Use plain text if no speaker labels formatted_text = transcript['text'] return {'formattedText': formatted_text} ``` This will output a `formattedText` field containing the transcript with speaker labels. ### Step 6: Save to Google Drive To save the formatted transcript back to Google Drive: n8n Convert to File node configured to output a text file Add a `Convert to File` node and configure it: - **Operation**: Convert to Text File - **Text Input Field**: formattedText - **Put Output File in Field**: data n8n Google Drive Upload file node configured to save the transcript Add a Google Drive `Upload file` node and configure it: - **Credential to connect with**: Select your Google Drive credential - **Resource**: File - **Operation**: Upload - **Input Data Field Name**: data - **File Name**: `{{ $('Get a transcription').item.json.id }}` (or drag and drop from input data) - **Parent Drive**: From list > My Drive - **Parent Folder**: From list > transcripts The formatted transcript will be saved as a `.txt` file in your specified Google Drive folder: Formatted transcript text file saved in the Google Drive transcripts folder ### Step 7: Delete Transcript (Optional) n8n AssemblyAI Delete a transcription node at the end of the workflow Once you're done processing the transcript, as it is now saved in Google Drive, you can **optionally** delete it from AssemblyAI's servers. Add an AssemblyAI `Delete a transcription` node at the end of your workflow. Configure the node parameters: - **Credential to connect with**: Select your AssemblyAI Account credential - **Resource**: Transcript - **Operation**: Delete - **Transcript ID**: `{{ $('Get a transcription').item.json.id }}` - You can drag and drop the `id` field from the `Get a transcription` node response into this field The transcript will be permanently deleted from AssemblyAI's servers. ### Step 8: Deploy Once the transcript has been deleted, the workflow is complete! At this point, it should look like this: Complete n8n AssemblyAI transcription workflow with all connected nodes Now you can deploy the workflow to n8n Cloud by navigating back to the n8n Cloud **Overview** page, locating your workflow, and moving the slider icon to the `Active` position. n8n Cloud Overview page with the workflow slider set to Active And we're done! Add a file to `/audio_files` in Google Drive, and within a few seconds to a few minutes (depending on the `Poll Times` set for the [Google Drive trigger](/integrations/n-8-n#choose-your-trigger) and the duration of the audio file), you should see the transcript appear in the `/transcripts` Google Drive folder. ## Conclusion In this tutorial, you built a complete AssemblyAI transcription workflow in n8n that automatically processes audio files from Google Drive, submits them to AssemblyAI for transcription, polls for completion, formats the transcript with speaker labels, saves the output back to Google Drive as a `.txt` file, and finally deletes the transcript from AssemblyAI's servers. This is just one simple idea, but the possibilities are endless! You can customize this workflow further by adding additional AssemblyAI features and products like sentiment analysis, entity detection, and LLM Gateway (see the **"What can you do with AssemblyAI?"** section of [this page](https://n8n.io/integrations/assemblyai/) for all available actions), or by integrating with other services like Slack, OpenAI, Supabase, and the hundreds of other [official n8n integrations](https://n8n.io/integrations/). ## Additional Resources - [AssemblyAI n8n Integration Page](https://n8n.io/integrations/assemblyai/) - [AssemblyAI n8n Integration GitHub Repo](https://github.com/gsharp-aai/n8n-nodes-assemblyai) ## Need some help? If you get stuck, think something is broken or missing from our n8n integration, or just have some questions, we'd love to help you out! Contact our support team directly via support@assemblyai.com or open a [support ticket](https://www.assemblyai.com/contact/support). --- # The Postman collection for the AssemblyAI API URL: https://www.assemblyai.com/docs/integrations/postman Source: docs/integrations/postman.mdx Navigation: Overview > Use cases & integrations > Integrations > No-code tools Description: Use the AssemblyAI API Postman collection to experiment with our API. Postman is a user-friendly tool for testing API endpoints. The AssemblyAI API Postman collection contains all the HTTP requests you can make to the AssemblyAI API, so you don't need to write them yourself. ## Quickstart Open the [AssemblyAI API collection](https://assembly.ai/postman) in Postman and click the **Fork** button. This will create a copy of the collection that you can edit. ![Click Fork on the AssemblyAI API collection](/assets/img/integrations/postman/1-fork-collection.png) Fill out the form and click **Fork Collection**. Next, click on the **Variables** tab and configure the `apiKey` variable with your AssemblyAI API key. You can find your AssemblyAI API key in the [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home). ![Configure AssemblyAI API key in Postman collection](/assets/img/integrations/postman/2-configure-variables.png) Let's upload an audio file: 1. Open the **Files > Upload a media file** request 2. Switch to the **Body** tab 3. Change the dropdown from **none** to **binary** 4. Select an audio file of your choosing, or [download this sample audio file](https://assembly.ai/nbc.mp3) 5. Click the **Send** button ![Send upload audio file to AssemblYAI API HTTP request](/assets/img/integrations/postman/3-upload-file-request.png) Inspect the **Body** of the response and copy the `upload_url` value. ![The response for uploading an audio file to AssemblYAI API](/assets/img/integrations/postman/4-upload-file-response.png) Now that the audio file is uploaded, you can transcribe the audio file. 1. Open the **Transcripts > Transcribe audio** request 2. Switch to the **Body** tab 3. Find the `audio_url` property and update it to the `upload_url` value from the previous request. 4. Remove all the other properties. Optionally, you can leave any property that you do want to use. 5. Click the **Send** button ![Create a transcript HTTP request](/assets/img/integrations/postman/5-create-transcript-request.png) Inspect the **Body** of the response and copy the `id` value. ![Create a transcript HTTP response](/assets/img/integrations/postman/6-create-transcript-response.png) The transcription job will take longer depending on the duration of the file. You need to check the `status` property to check if a transcript is ready. The `status` typically goes from `processing` to `completed`. The `status` can also be `queued` if the job is waiting to be processed (for example, when you've exceeded your rate limit), in which case it will move to `processing` once a slot is available. The `status` can also become `error` at any point. If an error occurs, you can find the error message under the `error` property. 1. Open the **Transcripts > Get transcript** request 2. Find the `transcript_id` under **Path variables** and update the value with the `id` value from the previous request. 3. Remove all the other properties. Optionally, you can leave any property that you do want to use. 4. Click the **Send** button ![Get a transcript HTTP request and response](/assets/img/integrations/postman/7-get-transcript.png) Inspect the **Body** of the response to check if the `status` is `completed` or `error`. If not, resend the request until it is `completed` or `error`. Now that you have a completed transcript, you can send these other HTTP requests with your current transcript ID: - Transcripts - Get subtitles for transcript - Get sentences in transcript - Get paragraphs in transcript - Search words in transcript - Get redacted audio (if PII audio redaction is enabled) --- # Build a Zoom Real-time transcription bot with Recall.ai URL: https://www.assemblyai.com/docs/integrations/recall Source: docs/integrations/recall.mdx Navigation: Overview > Use cases & integrations > Integrations > Meeting transcriber tools Description: Build a Zoom Real-time transcription bot with Recall.ai documentation. A real-time transcription bot that integrates [Recall.ai](https://www.recall.ai/assemblyai) with [AssemblyAI](https://www.assemblyai.com) to provide live transcription of Zoom meetings. ## Quickstart ```bash # 1. Clone and install git clone https://github.com/AssemblyAI/assemblyai-recallai-zoom-bot.git cd assemblyai-recallai-zoom-bot npm install # 2. Run ngrok and copy the ngrok URL # ngrok http 8000 # 3. Configure your .env file and edit it with your API keys and ngrok URL cp .env.example .env # 4a. Open a new terminal and run # node webhook.js # 4b. Open another terminal and run # node zoomBot.js ``` ## Prerequisites - [Node.js](https://nodejs.org/en/) (v14 or higher) - [ngrok](https://ngrok.com) - [Installation guide](https://ngrok.com/download) - Recall.ai API key with AssemblyAI configured ## Step-by-step Follow this step-by-step guide to set up and run the transcription bot. ### Step 1: Get your API keys **1.1 Choose your Recall.ai region and get your API Key:** | Region | Dashboard | RECALL_REGION | | ---------------- | ------------------------------------------------------------ | ---------------- | | US Pay-as-you-go | [us-west-2.recall.ai](https://us-west-2.recall.ai) | `us-west-2` | | US Monthly plan | [us-east-1.recall.ai](https://us-east-1.recall.ai) | `us-east-1` | | EU | [eu-central-1.recall.ai](https://eu-central-1.recall.ai) | `eu-central-1` | | Japan | [ap-northeast-1.recall.ai](https://ap-northeast-1.recall.ai) | `ap-northeast-1` | **1.2 Configure AssemblyAI in your Recall.ai dashboard:** - Get an AssemblyAI API key from [assemblyai.com](https://www.assemblyai.com) - In your Recall.ai dashboard (same region as step 1.1), navigate to the transcription providers section - Add your AssemblyAI API key to enable AssemblyAI as a transcript provider This step is required for the bot to work with AssemblyAI transcription! ### Step 2: Clone the example repo and install dependencies Open a new terminal and run the following commands: ```bash git clone https://github.com/AssemblyAI/assemblyai-recallai-zoom-bot.git cd assemblyai-recallai-zoom-bot npm install ``` ### Step 3: Configure environment Set up your environment variables using the example `.env` file in the repo: ```bash cp .env.example .env ``` Edit `.env` with your Recall values. ```bash RECALL_API_KEY=your_recall_api_key RECALL_REGION=us-west-2 ``` Set `RECALL_REGION` to according to your Recall region in Step 1.1 ### Step 4: Start ngrok tunnel In a new terminal, start ngrok to create a public URL for your webhook: ```bash ngrok http 8000 ``` The output should look like this: ``` ngrok by @inconshreveable Session Status online Account your-account@email.com Version 2.3.40 Region United States (us) Web Interface http://127.0.0.1:4040 Forwarding https://abc123.ngrok.io -> http://localhost:8000 Forwarding http://abc123.ngrok.io -> http://localhost:8000 ``` Copy the https URL (e.g., `https://abc123.ngrok.io`) - you'll need this for the next step. ### Step 5: Update Webhook URL Edit your `.env` file and set `WEBHOOK_URL` to your ngrok URL: ```bash WEBHOOK_URL=https://abc123.ngrok.io ``` Use the exact URL from ngrok output (no trailing slash) ### Step 6: Start webhook server In another terminal, start the webhook server (receives transcripts): ```bash node webhook.js ``` ### Step 7: Start the bot CLI In another terminal, start the bot CLI (manages meeting connection): ```bash node zoomBot.js ``` Once you run this command, you'll see this output: ``` [BOT] Starting Recall.ai bot for Zoom → AssemblyAI integration [BOT] Configured for region: us-west-2 [BOT] API endpoint: https://us-west-2.recall.ai/api/v1 [BOT] Starting application... [BOT] Validating configuration... [BOT] ✓ Configuration valid [BOT] ✓ Region: us-west-2 [BOT] ✓ Webhook: https://abc123.ngrok.io [BOT] Ready to join meeting and start transcription What is your meeting URL?: ``` ### Step 8: Join meeting and start transcription Enter your Zoom meeting URL when prompted: ``` What is your meeting URL?: https://zoom.us/j/123456789 ``` And your terminal will look like this: ``` [BOT] Creating bot for meeting: https://zoom.us/j/123456789 [BOT] Webhook endpoint: https://abc123.ngrok.io/meeting_transcript [API] POST /bot - Creating bot with AssemblyAI integration [API] Bot created successfully [BOT] Bot ID: bot_abc123 [BOT] Status: joining_call [BOT] ✓ Bot deployed successfully [BOT] Bot is joining meeting and will start sending transcripts to webhook [BOT] Check Terminal 2 for real-time transcripts Type "STOP" to end transcription: ``` In your second terminal (running `webhook.js`), you should see a meeting transcript of your participants: ``` [WEBHOOK] Server running on port 8000 [WEBHOOK] Ready to receive transcripts from Recall.ai [WEBHOOK] Integration: Zoom → Recall.ai → AssemblyAI → This webhook [TRANSCRIPT] PARTIAL - John Doe: Hello everyone [TRANSCRIPT] PARTIAL - John Doe: Hello everyone, welcome to [TRANSCRIPT] FINAL - John Doe: Hello everyone, welcome to today's meeting. [TRANSCRIPT] FINAL - Jane Smith: Thanks for joining, let's get started with the agenda. ``` ### Step 9: Stop transcription To stop transcribing, type "STOP" to end transcription on your third terminal running `node zoomBot.js`: ``` Type "STOP" to end transcription: STOP ``` ## Troubleshooting ### Environment Variable Errors #### `RECALL_API_KEY not found in .env file` Solution: Make sure you copied `.env.example` to `.env` and added your API key. #### `RECALL_REGION must be one of: us-west-2, us-east-1, eu-central-1, ap-northeast-1` Solution: Check your `.env` file and ensure `RECALL_REGION` matches where you got your API key ### Bot Creation Errors #### AssemblyAI not configured ```bash Failed to create bot: { recording_config: { transcript: { provider: [Object] } } } ``` Solution: Configure AssemblyAI in your Recall.ai dashboard (Step 1.2) ### Webhook Issues #### No transcripts appearing ```bash [WEBHOOK] Server running on port 8000 [WEBHOOK] Ready to receive transcripts from Recall.ai [WEBHOOK] Integration: Zoom → Recall.ai → AssemblyAI → This webhook (no transcript output) ``` Solution: - Verify `WEBHOOK_URL` in `.env` matches your ngrok URL exactly - Ensure ngrok is still running (it may timeout after inactivity) - Check that the webhook server was started before the bot #### ngrok connection refused ```bash Failed to complete tunnel connection ``` Solution: - Restart ngrok: `ngrok http 8000` - Update `WEBHOOK_URL` in `.env` with the new ngrok URL - Ensure port 8000 is available ### Common network issues #### Timeout connecting to Recall.ai ```bash timeout of 10000ms exceeded ``` Solution: - Check your internet connection - Verify your `RECALL_REGION` is correct - Try again after a few seconds #### Bot appears to join but no transcripts Solution: - Ensure people are speaking in the meeting - Check that meeting participants have unmuted their microphones - Verify AssemblyAI is properly configured in Recall.ai dashboard **Stuck?** Contact our support team at support@assemblyai.com or create a [support ticket](https://www.assemblyai.com/contact/support). --- # Transcribe Your Zoom Meetings URL: https://www.assemblyai.com/docs/integrations/zoom-rtms Source: docs/integrations/zoom-rtms.mdx Navigation: Overview > Use cases & integrations > Integrations > Meeting transcriber tools Description: Transcribe Your Zoom Meetings documentation. This guide creates a Node.js service that captures audio from Zoom Real-Time Media Streams (RTMS) and provides both real-time and asynchronous transcription using AssemblyAI. **Zoom RTMS Documentation** For complete Zoom RTMS documentation, visit https://developers.zoom.us/docs/rtms/ ## Features - **Real-time Transcription**: Live transcription during meetings using AssemblyAI's streaming API - **Asynchronous Transcription**: Complete post-meeting transcription with advanced features - **Flexible Audio Modes**: - Mixed stream (all participants combined) - Individual participant streams transcribed - **Multichannel Audio Support**: Separate channels for different participants - **Configurable Processing**: Enable/disable real-time or async transcription independently ## Setup ### Prerequisites - Node.js 16+ - FFmpeg installed on your system - Zoom RTMS Developer Preview access - AssemblyAI API key - ngrok (for local development and testing) ### Installation 1. **Clone the example repository and install dependencies**: ```bash git clone https://github.com/zkleb-aai/assemblyai-zoom-rtms.git cd assemblyai-zoom-rtms npm install ``` 2. **Configure environment variables**: ```bash cp .env.example .env ``` Fill in your `.env` file: ```env # Zoom Configuration ZM_CLIENT_ID=your_zoom_client_id ZM_CLIENT_SECRET=your_zoom_client_secret ZOOM_SECRET_TOKEN=your_webhook_secret_token # AssemblyAI Configuration ASSEMBLYAI_API_KEY=your_assemblyai_api_key # Service Configuration PORT=8080 REALTIME_ENABLED=true REALTIME_MODE=mixed ASYNC_ENABLED=true AUDIO_CHANNELS=mono AUDIO_SAMPLE_RATE=16000 TARGET_CHUNK_DURATION_MS=100 ``` ### Local development with ngrok For testing and development, you can use ngrok to expose your local server to the internet: 1. **Install ngrok**: Download from [ngrok.com](https://ngrok.com/) or install via package manager: ```bash # macOS brew install ngrok # Windows (chocolatey) choco install ngrok # Or download directly from ngrok.com ``` 2. **Start your local server**: ```bash npm start ``` 3. **In a separate terminal, start ngrok**: ```bash ngrok http 8080 ``` 4. **Copy the ngrok URL**: ngrok will display a forwarding URL like: ``` Forwarding https://example-abc123.ngrok-free.app -> http://localhost:8080 ``` 5. **Use the ngrok URL in your Zoom app webhook configuration**: ``` https://example-abc123.ngrok-free.app/webhook ``` ### Configuration options #### Real-time transcription - `REALTIME_ENABLED`: Enable/disable live transcription (default: `true`) - `REALTIME_MODE`: - `mixed`: Single stream with all participants combined - `individual`: Separate streams per participant #### Audio settings - `AUDIO_CHANNELS`: `mono` or `multichannel` - `AUDIO_SAMPLE_RATE`: Audio sample rate in Hz (default: `16000`) - `TARGET_CHUNK_DURATION_MS`: Audio chunk duration for streaming (default: `100`) #### Async transcription - `ASYNC_ENABLED`: Enable/disable post-meeting transcription (default: `true`) ## Usage ### Start the service ```bash npm start ``` The service will start on the configured port (default: 8080) and display: ``` 🎧 Zoom RTMS to AssemblyAI Transcription Service 📋 Configuration: Real-time: ✅ (mixed) Audio: mono @ 16000Hz Async: ✅ 🚀 Server running on port 8080 📡 Webhook endpoint: http://localhost:8080/webhook ``` ### Configure Zoom webhook 1. In your Zoom App configuration, set the webhook endpoint to: ``` # For production https://your-domain.com/webhook # For local development with ngrok https://example-abc123.ngrok-free.app/webhook ``` 2. Subscribe to these events: - `meeting.rtms_started` - `meeting.rtms_stopped` ### Testing with ngrok When using ngrok for testing: 1. **Keep ngrok running**: The ngrok tunnel must remain active during testing 2. **Update webhook URL**: If you restart ngrok, you'll get a new URL that needs to be updated in your Zoom app configuration 3. **Monitor ngrok logs**: ngrok shows incoming webhook requests in its terminal output 4. **Free tier limitations**: The free ngrok tier has some limitations; consider upgrading for heavy testing ### Real-time output During meetings, you'll see live transcription: ``` 🚀 AssemblyAI session started: [abc12345] 🎙️ [abc12345] Hello everyone, welcome to the meeting 📝 [abc12345] FINAL: Hello everyone, welcome to the meeting. ``` ### Post-meeting files After each meeting, the service generates: - `transcript_[meeting_uuid].json` - Full AssemblyAI response with metadata - `transcript_[meeting_uuid].txt` - Plain text transcript ## Advanced configuration ### AssemblyAI features Modify the `ASYNC_CONFIG` object in the code to enable additional features: ```javascript const ASYNC_CONFIG = { speaker_labels: true, // Speaker identification auto_chapters: true, // Automatic chapter detection sentiment_analysis: true, // Sentiment analysis entity_detection: true, // Named entity recognition redact_pii: true, // PII redaction summarization: true, // Auto-summarization auto_highlights: true, // Key highlights }; ``` See [AssemblyAI's API documentation](/api-reference/transcripts/submit) for all available options. ### Audio processing modes #### Mixed mode (default) - Single audio stream combining all participants - Most efficient for general transcription - Best for meetings with clear speakers #### Individual mode - Separate transcription stream per participant - Better speaker attribution - Higher resource usage #### Multichannel audio - Separate audio channels for different participants - Enables advanced speaker separation - Requires `AUDIO_CHANNELS=multichannel` ## API endpoints ### `POST` /webhook Handles Zoom RTMS webhook events: - URL validation - Meeting start/stop events - Automatic RTMS connection setup ## Error handling The service includes comprehensive error handling: - Automatic reconnection for dropped connections - Graceful cleanup on meeting end - Audio buffer flushing to prevent data loss - Temporary file cleanup ## Monitoring ### Real-time logs - Connection status updates - Audio processing statistics - Transcription progress - Error notifications ### Example log output ``` 📡 Connecting to Zoom signaling for meeting abc123 ✅ Zoom signaling connected for meeting abc123 🎵 Connecting to Zoom media for meeting abc123 ✅ Zoom media connected for meeting abc123 🚀 Started audio streaming for meeting abc123 🎵 [abc12345] 100 chunks, 32768 bytes, 10.2s 📝 [abc12345] FINAL: This is the final transcription. ``` ### Development workflow 1. Start your local server: `npm start` 2. Start ngrok in another terminal: `ngrok http 8080` 3. Update your Zoom app webhook URL with the ngrok URL 4. Test with Zoom meetings 5. Monitor logs in both your app and ngrok terminals --- # Integrate Telnyx with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/telnyx Source: docs/integrations/telnyx.mdx Navigation: Overview > Use cases & integrations > Integrations > Telephony tools Description: Build voice agents with Telnyx and AssemblyAI using Pipecat or LiveKit. [Telnyx](https://telnyx.com/) is a global connectivity platform that provides programmable voice, messaging, and wireless services. By combining Telnyx with AssemblyAI, you can build real-time voice agents with industry-leading speech recognition accuracy and advanced turn detection. This guide shows you how to integrate Telnyx with AssemblyAI using two popular voice agent orchestrators: LiveKit and Pipecat. ## LiveKit Telephony Integration [LiveKit](https://livekit.io/) is a real-time communication platform for building voice, video, and data applications. You can integrate Telnyx SIP trunking with LiveKit to enable phone calls to your voice agents that use AssemblyAI for speech recognition. ### How it works Telnyx SIP trunks bridge phone calls into LiveKit rooms as special SIP participants. Your LiveKit agent (configured with AssemblyAI STT) connects to the room and interacts with the caller. The flow is: 1. **Phone call** → Telnyx SIP trunk 2. **Telnyx** → LiveKit room (creates SIP participant) 3. **LiveKit agent** (with AssemblyAI STT) → joins room and handles conversation ### Before you begin - **Telnyx account**: Create an account and purchase a phone number at [telnyx.com](https://telnyx.com) - **LiveKit project**: Get your SIP URI from your [LiveKit project settings](https://cloud.livekit.io) - **LiveKit CLI**: Install the [LiveKit CLI](https://docs.livekit.io/home/cli/cli-setup/) on your computer - **Environment variables**: Configure `LIVEKIT_URL`, `LIVEKIT_API_KEY`, `LIVEKIT_API_SECRET`, and `ASSEMBLYAI_API_KEY` ### Step 1: Configure Telnyx SIP connection Configure your Telnyx SIP connection to route calls to LiveKit. Follow the detailed setup guide: [Telnyx LiveKit SIP Configuration Guide](https://developers.telnyx.com/docs/voice/sip-trunking/livekit-configuration-guide) Key steps: 1. Create a SIP connection in Telnyx Mission Control Portal 2. Set connection type to FQDN and provide your LiveKit SIP URI 3. Configure outbound call authentication (username/password) 4. Assign your phone number(s) to the SIP connection ### Step 2: Configure LiveKit SIP trunks Create inbound and outbound SIP trunks in LiveKit using the LiveKit CLI. #### Inbound trunk Create `inboundTrunk.json`: ```json { "trunk": { "name": "Telnyx Inbound Trunk", "numbers": ["YOUR_TELNYX_NUMBER"] } } ``` Create the trunk: ```bash lk sip inbound create inboundTrunk.json ``` Save the returned `SIPTrunkID` for the next step. #### Dispatch rule Create `dispatchRule.json` to route incoming calls to your agent: ```json { "name": "Agent Dispatch Rule", "trunk_ids": ["YOUR_TRUNK_ID"], "rule": { "dispatchRuleIndividual": { "roomPrefix": "call-" } }, "roomConfig": { "agents": [ { "agentName": "my-telephony-agent" } ] } } ``` Create the dispatch rule: ```bash lk sip dispatch create dispatchRule.json ``` This automatically dispatches your agent to each incoming call in a new room. #### Outbound trunk (optional) For outbound calling, create `outboundTrunk.json`: ```json { "trunk": { "name": "Telnyx Outbound Trunk", "address": "sip.telnyx.com", "numbers": ["YOUR_TELNYX_NUMBER"], "auth_username": "YOUR_OUTBOUND_USER", "auth_password": "YOUR_OUTBOUND_PASS" } } ``` Create the trunk: ```bash lk sip outbound create outboundTrunk.json ``` ### Step 3: Build your LiveKit agent with AssemblyAI Create a LiveKit agent that uses AssemblyAI for speech recognition. Once your SIP trunks and dispatch rules are configured, no special telephony code is required—the agent simply joins the room when a call comes in. #### Prerequisites Setup and activate a virtual environment: ```bash python -m venv venv source venv/bin/activate ``` Install the required dependencies: ```bash pip install livekit-agents livekit-plugins-assemblyai livekit-plugins-openai livekit-plugins-silero livekit-plugins-rime ``` Download model files: ```bash python agent.py download-files ``` #### Agent code Here's a minimal agent using AssemblyAI STT: ```python expandable import asyncio from dotenv import load_dotenv from livekit import rtc from livekit.agents import AutoSubscribe, JobContext, WorkerOptions, cli, llm from livekit.agents.voice_assistant import VoiceAssistant from livekit.plugins import assemblyai, openai, silero, rime load_dotenv() async def entrypoint(ctx: JobContext): initial_ctx = llm.ChatContext().append( role="system", text=( "You are a helpful voice assistant. Your interface with users will be voice. " "Use short and concise responses, avoiding unpronounceable punctuation." ), ) await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY) assistant = VoiceAssistant( vad=silero.VAD.load(), stt=assemblyai.STT(), llm=openai.LLM(), tts=rime.TTS(), chat_ctx=initial_ctx, ) assistant.start(ctx.room) await asyncio.sleep(1) await assistant.say("Hello! How can I help you today?", allow_interruptions=True) if __name__ == "__main__": cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint)) ``` #### Environment variables Set the following environment variables: ```bash export LIVEKIT_URL="wss://your-livekit-server" export LIVEKIT_API_KEY="YOUR_LIVEKIT_API_KEY" export LIVEKIT_API_SECRET="YOUR_LIVEKIT_API_SECRET" export ASSEMBLYAI_API_KEY="YOUR_ASSEMBLYAI_API_KEY" export OPENAI_API_KEY="YOUR_OPENAI_API_KEY" export RIME_API_KEY="YOUR_RIME_API_KEY" ``` #### Run the agent Start your agent: ```bash python agent.py dev ``` Now dial your Telnyx phone number to test the integration. The call will be routed through Telnyx → LiveKit → your agent with AssemblyAI speech recognition. ### Testing outbound calls To test an outbound call, create `sipParticipant.json`: ```json { "sip_trunk_id": "YOUR_OUTBOUND_TRUNK_ID", "sip_call_to": "+15105551234", "room_name": "test-outbound-call", "participant_identity": "outbound-test", "participant_name": "Test Call" } ``` Place the call: ```bash lk sip participant create sipParticipant.json ``` ### Troubleshooting **Call connects but agent never speaks**: Verify your dispatch rule includes `roomConfig.agents` with the correct `agentName`. Without this, the agent won't be automatically dispatched to the room. **Audio quality issues**: Check your Telnyx SIP connection settings and ensure proper codec configuration. See the [Telnyx troubleshooting guide](https://developers.telnyx.com/docs/voice/sip-trunking/livekit-configuration-guide#troubleshooting). ## Pipecat Integration [Pipecat](https://github.com/pipecat-ai/pipecat) is an open-source framework for building voice and multimodal conversational AI agents. You can integrate Telnyx Media Streaming with Pipecat to enable phone calls to your voice agents that use AssemblyAI for speech recognition. This guide covers both dial-in (users call your number) and dial-out (your bot calls users) functionality. ### Prerequisites - **Telnyx account**: Create an account and purchase phone numbers at [telnyx.com](https://telnyx.com) - **Public server or tunnel**: For dial-out, you'll need a public-facing server or tunneling service like [ngrok](https://ngrok.com/) - **API keys**: Get API keys for AssemblyAI, OpenAI (or other LLM), and your preferred TTS service Install the required dependencies: ```bash pip install pipecat-ai python-dotenv loguru ``` Set up your environment variables: ```bash TELNYX_API_KEY=your_telnyx_api_key ASSEMBLYAI_API_KEY=your_assemblyai_api_key OPENAI_API_KEY=your_openai_api_key DEEPGRAM_API_KEY=your_deepgram_api_key # or other TTS provider ``` ### How it works **Dial-in flow**: 1. User calls your Telnyx number 2. Telnyx executes your TeXML application which establishes a WebSocket connection 3. Telnyx opens a WebSocket to your server with real-time audio and call metadata 4. Your bot processes the audio using the Pipecat pipeline with AssemblyAI STT 5. The bot responds with audio sent back to Telnyx over WebSocket 6. Telnyx plays the audio to the caller in real-time **Dial-out flow**: 1. Your application triggers a dial-out via API 2. Server initiates a Telnyx call using the Call Control API 3. Telnyx establishes the call and opens a WebSocket connection 4. Your bot joins the WebSocket and sets up the pipeline 5. The recipient answers and is connected to your bot 6. The bot handles the conversation with real-time audio streaming ### Dial-in Setup #### Step 1: Create TeXML application Telnyx uses TeXML (Telnyx Extensible Markup Language) to control call flow. Create a TeXML application that establishes a WebSocket connection to your bot: ```xml ``` The `bidirectionalMode="rtp"` parameter enables real-time audio streaming in both directions. **Custom data with query parameters**: You can pass custom data to your bot by adding query parameters to the WebSocket URL: ```xml ``` #### Step 2: Configure Telnyx phone number 1. Go to the [Telnyx Portal](https://portal.telnyx.com) 2. Navigate to Voice → Programmable Voice → TeXML Applications 3. Create a new TeXML Application with your WebSocket URL 4. Assign the TeXML Application to your phone number #### Step 3: Build your bot Here's a complete dial-in bot using AssemblyAI for speech recognition: ```python expandable import os from dotenv import load_dotenv from loguru import logger from pipecat.audio.vad.silero import SileroVADAnalyzer from pipecat.frames.frames import EndFrame from pipecat.pipeline.pipeline import Pipeline from pipecat.pipeline.runner import PipelineRunner from pipecat.pipeline.task import PipelineParams, PipelineTask from pipecat.processors.aggregators.openai_llm_context import OpenAILLMContext from pipecat.runner.utils import parse_telephony_websocket from pipecat.serializers.telnyx import TelnyxFrameSerializer from pipecat.services.assemblyai.stt import AssemblyAISTTService from pipecat.services.openai.llm import OpenAILLMService from pipecat.services.deepgram.tts import DeepgramTTSService from pipecat.transports.websocket.fastapi import ( FastAPIWebsocketTransport, FastAPIWebsocketParams, ) load_dotenv() async def run_bot(websocket): """Run the voice agent bot with AssemblyAI STT.""" # Parse Telnyx WebSocket data - automatically extracts call information transport_type, call_data = await parse_telephony_websocket(websocket) # Extract call information (automatically provided by Telnyx) stream_id = call_data["stream_id"] call_control_id = call_data["call_control_id"] outbound_encoding = call_data["outbound_encoding"] from_number = call_data["from"] # Caller's number to_number = call_data["to"] # Your Telnyx number logger.info(f"Incoming call from {from_number} to {to_number}") # Create Telnyx serializer with call details serializer = TelnyxFrameSerializer( stream_id=stream_id, call_control_id=call_control_id, api_key=os.getenv("TELNYX_API_KEY"), ) # Configure WebSocket transport transport = FastAPIWebsocketTransport( websocket=websocket, params=FastAPIWebsocketParams( audio_in_enabled=True, audio_out_enabled=True, add_wav_header=False, vad_analyzer=SileroVADAnalyzer(), serializer=serializer, ), ) # Configure AI services stt = AssemblyAISTTService(api_key=os.getenv("ASSEMBLYAI_API_KEY")) llm = OpenAILLMService(api_key=os.getenv("OPENAI_API_KEY"), model="gpt-5-mini") tts = DeepgramTTSService(api_key=os.getenv("DEEPGRAM_API_KEY")) # Customize bot behavior based on call information messages = [ { "role": "system", "content": f"You are a helpful voice assistant. The caller is calling from {from_number}. Keep responses concise and conversational." } ] context = OpenAILLMContext(messages) context_aggregator = llm.create_context_aggregator(context) # Build pipeline pipeline = Pipeline([ transport.input(), stt, context_aggregator.user(), llm, tts, transport.output(), context_aggregator.assistant(), ]) task = PipelineTask( pipeline, params=PipelineParams( allow_interruptions=True, audio_in_sample_rate=8000, audio_out_sample_rate=8000, ), ) @transport.event_handler("on_client_disconnected") async def on_client_disconnected(transport, client): logger.info("Call ended") await task.queue_frame(EndFrame()) # Run the pipeline runner = PipelineRunner() await runner.run(task) ``` See the [complete dial-in example](https://github.com/pipecat-ai/pipecat-examples/tree/main/telnyx-chatbot/inbound) for full implementation details including server setup. ### Dial-out Setup Dial-out allows your bot to initiate calls to phone numbers using Telnyx's outbound calling capabilities. #### How dial-out works 1. Your application triggers a dial-out (via API call or user action) 2. Server initiates a Telnyx call using the Call Control API 3. Telnyx establishes the call and opens a WebSocket connection 4. Your bot joins the WebSocket and sets up the pipeline 5. The recipient answers and is connected to your bot 6. The bot handles the conversation with real-time audio streaming #### Bot configuration for dial-out The dial-out bot configuration is similar to dial-in. Telnyx automatically provides call information in the WebSocket messages: ```python # Parse WebSocket data (same as dial-in) transport_type, call_data = await parse_telephony_websocket(websocket) # Extract call information stream_id = call_data["stream_id"] call_control_id = call_data["call_control_id"] from_number = call_data["from"] # Your Telnyx number to_number = call_data["to"] # Target number you're calling # Customize bot behavior for outbound calls greeting = f"Hi! This is an automated call from {from_number}. How are you today?" ``` The transport and pipeline configuration are identical to dial-in. See the [complete dial-out example](https://github.com/pipecat-ai/pipecat-examples/tree/main/telnyx-chatbot/outbound) for full server implementation with outbound call creation. ### Key Features **Audio format**: Telnyx Media Streaming uses 8kHz mono audio with 16-bit PCM encoding. Configure your pipeline accordingly: ```python task = PipelineTask( pipeline, params=PipelineParams( audio_in_sample_rate=8000, audio_out_sample_rate=8000, allow_interruptions=True, ), ) ``` **Automatic call termination**: When you provide Telnyx API credentials to the `TelnyxFrameSerializer`, it automatically ends calls when your pipeline ends: ```python serializer = TelnyxFrameSerializer( stream_id=stream_id, call_control_id=call_control_id, api_key=os.getenv("TELNYX_API_KEY"), # Enables auto-termination ) ``` **Built-in call information**: Unlike other providers, Telnyx automatically includes caller information (to/from numbers) in the WebSocket messages, eliminating the need for custom webhook servers in basic dial-in scenarios. ### Advanced Configuration For production use cases, you can customize AssemblyAI's turn detection and add keyterms for improved accuracy: ```python from pipecat.services.assemblyai.stt import AssemblyAISTTService, AssemblyAIConnectionParams # Configure AssemblyAI with custom parameters stt_params = AssemblyAIConnectionParams( sample_rate=8000, end_of_turn_confidence_threshold=0.4, min_turn_silence=400, max_turn_silence=1280, keyterms_prompt=["NPI", "TIN", "CMS", "PTAN", "CPT", "CDT", "DOB", "SSN"] ) stt = AssemblyAISTTService( api_key=os.getenv("ASSEMBLYAI_API_KEY"), connection_params=stt_params, ) ``` **Turn detection options**: AssemblyAI has built-in VAD and turn detection. You can either: - Use AssemblyAI's turn detection (recommended for best accuracy) - Use Silero VAD by including `vad_analyzer=SileroVADAnalyzer()` in the transport params **Keyterms**: Add domain-specific terms to improve recognition accuracy for specialized vocabulary. ## Resources - [LiveKit Documentation](https://docs.livekit.io/) - [LiveKit Telephony Quickstart](https://docs.livekit.io/agents/start/telephony/) - [Telnyx SIP Trunk Configuration Guide](https://developers.telnyx.com/docs/voice/sip-trunking/livekit-configuration-guide) - [Telnyx LiveKit SIP Setup](https://developers.telnyx.com/docs/voice/sip-trunking/livekit-configuration-guide) - [Pipecat Documentation](https://docs.pipecat.ai/) - [Pipecat Telnyx WebSocket Guide](https://docs.pipecat.ai/guides/telephony/telnyx-websockets) - [Pipecat Telnyx Examples](https://github.com/pipecat-ai/pipecat-examples/tree/main/telnyx-chatbot) --- # Integrate Twilio with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/twilio Source: docs/integrations/twilio.mdx Navigation: Overview > Use cases & integrations > Integrations > Telephony tools Description: Transcribe Twilio Voice data using AssemblyAI. Twilio is a programmable communication platform for voice, messaging, and email. By combining Twilio with AssemblyAI, you can transcribe voice calls in [real-time](/streaming/getting-started/transcribe-streaming-audio), and voice recordings and voice messages [asynchronously](/pre-recorded-audio). Combine transcription with our [speech understanding models](/speech-understanding/summarization) to analyze the calls and messages. Blog posts: - [Transcribe a phone call in real-time using Python with AssemblyAI and Twilio](https://www.assemblyai.com/blog/transcribe-phone-call-real-time-python/) - [Transcribe Phone Calls in Real-Time using Node.js with AssemblyAI, and Twilio](https://www.twilio.com/en-us/blog/phone-call-transcription-assemblyai-twilio-node) - [Transcribe phone calls in real-time using C# .NET with AssemblyAI and Twilio](https://www.twilio.com/en-us/blog/transcribe-phone-calls-real-time-csharp-assemblyai-twilio) - [Transcribe phone calls in real-time in Go with Twilio and AssemblyAI](https://www.assemblyai.com/blog/transcribe-phone-calls-in-realtime-in-go-with-twilio-and-assemblyai/) - [Answer Questions about Twilio Voice Recordings with AssemblyAI and LangChain.js](https://www.twilio.com/en-us/blog/qa-voice-recordings-assemblyai-langchain-js) - [Using Django & AssemblyAI for More Accurate Twilio Call Transcriptions](https://www.fullstackpython.com/blog/django-accurate-twilio-voice-transcriptions.html) Videos: - [Transcribe Twilio Phone Calls in Real-Time with AssemblyAI | JavaScript WebSockets Tutorial](https://youtu.be/3XmtJgWcOT0) --- # Transcribe Your Amazon Connect Recordings URL: https://www.assemblyai.com/docs/integrations/amazon-connect Source: docs/integrations/amazon-connect.mdx Navigation: Overview > Use cases & integrations > Integrations > Telephony tools Description: Transcribe Your Amazon Connect Recordings documentation. This guide walks through the process of setting up a transcription pipeline for Amazon Connect recordings using AssemblyAI. ## Get Started Before we begin, make sure you have: - An AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. - An AWS account. - An [Amazon Connect instance](https://docs.aws.amazon.com/connect/latest/adminguide/amazon-connect-instances.html). ## Step-by-Step Instructions In the AWS console, navigate to the **Amazon Connect** services page. Select your instance and then click into the **Data Storage** section. On this page, find the subsection named **Call Recordings** and note the S3 bucket path where your call recordings are stored, you'll need this for later. Amazon Connect Data Storage page showing the Call Recordings S3 bucket path Navigate to the **Lambda** services page, and create a new function. Set the runtime to **Python 3.13**. In the **Change default execution role** section, choose the option to create a new role with basic Lambda permissions. Assign a function name and then click **Create function**. AWS Lambda create function page with the Python 3.13 runtime selected In this new function, scroll down to the **Code Source** section and paste the following code into `lambda_function.py`. ```python expandable import json import os import boto3 import http.client import time from urllib.parse import unquote_plus import logging # Configure logging logger = logging.getLogger() logger.setLevel(logging.INFO) # Configuration settings # See config parameters here: https://www.assemblyai.com/docs/api-reference/transcripts/submit ASSEMBLYAI_CONFIG = { # 'language_code': 'en_us', # 'multichannel': True, # 'redact_pii': True, } # Initialize AWS services s3_client = boto3.client('s3') def get_presigned_url(bucket, key, expiration=3600): """Generate a presigned URL for the S3 object""" logger.info({ "message": "Generating presigned URL", "bucket": bucket, "key": key, "expiration": expiration }) s3_client_with_config = boto3.client( 's3', config=boto3.session.Config(signature_version='s3v4') ) return s3_client_with_config.generate_presigned_url( 'get_object', Params={'Bucket': bucket, 'Key': key}, ExpiresIn=expiration ) def delete_transcript_from_assemblyai(transcript_id, api_key): """ Delete transcript data from AssemblyAI's database using their DELETE endpoint. Args: transcript_id (str): The AssemblyAI transcript ID to delete api_key (str): The AssemblyAI API key Returns: bool: True if deletion was successful, False otherwise """ headers = { "authorization": api_key, "content-type": "application/json" } conn = http.client.HTTPSConnection("api.assemblyai.com") try: # Send DELETE request to AssemblyAI API conn.request("DELETE", f"/v2/transcript/{transcript_id}", headers=headers) response = conn.getresponse() # Check if deletion was successful (HTTP 200) if response.status == 200: response_data = json.loads(response.read().decode()) logger.info(f"Successfully deleted transcript {transcript_id} from AssemblyAI") return True else: error_message = response.read().decode() logger.error(f"Failed to delete transcript {transcript_id}: HTTP {response.status} - {error_message}") return False except Exception as e: logger.info(f"Error deleting transcript {transcript_id}: {str(e)}") return False finally: conn.close() def transcribe_audio(audio_url, api_key): """Transcribe audio using AssemblyAI API with http.client""" logger.info({"message": "Starting audio transcription"}) headers = { "authorization": api_key, "content-type": "application/json" } conn = http.client.HTTPSConnection("api.assemblyai.com") # Submit the audio file for transcription with config parameters request_data = {"audio_url": audio_url} # Add all configuration settings request_data.update(ASSEMBLYAI_CONFIG) json_data = json.dumps(request_data) conn.request("POST", "/v2/transcript", json_data, headers) response = conn.getresponse() if response.status != 200: raise Exception(f"Failed to submit audio for transcription: {response.read().decode()}") response_data = json.loads(response.read().decode()) transcript_id = response_data['id'] logger.info({"message": "Audio submitted for transcription", "transcript_id": transcript_id}) # Poll for transcription completion while True: conn = http.client.HTTPSConnection("api.assemblyai.com") conn.request("GET", f"/v2/transcript/{transcript_id}", headers=headers) polling_response = conn.getresponse() polling_data = json.loads(polling_response.read().decode()) if polling_data['status'] == 'completed': conn.close() logger.info({"message": "Transcription completed successfully"}) return polling_data # Return full JSON response instead of just text elif polling_data['status'] == 'error': conn.close() raise Exception(f"Transcription failed: {polling_data['error']}") conn.close() time.sleep(3) def lambda_handler(event, context): """Lambda function to handle S3 events and process audio files""" try: # Get the AssemblyAI API key from environment variables api_key = os.environ.get('ASSEMBLYAI_API_KEY') if not api_key: raise ValueError("ASSEMBLYAI_API_KEY environment variable is not set") # Process each record in the S3 event for record in event.get('Records', []): # Get the S3 bucket and key bucket = record['s3']['bucket']['name'] key = unquote_plus(record['s3']['object']['key']) # Generate a presigned URL for the audio file audio_url = get_presigned_url(bucket, key) # Get the full transcript JSON from AssemblyAI transcript_data = transcribe_audio(audio_url, api_key) # Prepare the transcript key - maintaining path structure but changing directory and extension transcript_key = key.replace('/CallRecordings/', '/AssemblyAITranscripts/', 1).replace('.wav', '.json') # Convert the JSON data to a string transcript_json_str = json.dumps(transcript_data, indent=2) # Upload the transcript JSON to the same bucket but in transcripts directory s3_client.put_object( Bucket=bucket, # Use the same bucket Key=transcript_key, Body=transcript_json_str, ContentType='application/json' ) logger.info({"message": "Transcript uploaded to transcript bucket successfully.", "key": transcript_key}) # Uncomment the following line to delete transcript data from AssemblyAI after saving to S3 # https://www.assemblyai.com/docs/api-reference/transcripts/delete # delete_transcript_from_assemblyai(transcript_data['id'], api_key) return { "statusCode": 200, "body": json.dumps({ "message": "Audio file(s) processed successfully", "detail": "Transcripts have been stored in the AssemblyAITranscripts directory" }) } except Exception as e: print(f"Error: {str(e)}") return { "statusCode": 500, "body": json.dumps({ "message": "Error processing audio file(s)", "error": str(e) }) } ``` At the top of the lambda function, you can edit the config to enable features for your transcripts. To see all available parameters, check out our [API reference](/api-reference/transcripts/submit). ```python ASSEMBLYAI_CONFIG = { # 'language_code': 'en_us', # 'multichannel': True, # 'redact_pii': True, } ``` If you would like to delete transcripts from AssemblyAI after completion, you can uncomment line **166** to enable the `delete_transcript_from_assemblyai` function. This ensures the transcript data is only saved on your S3 database and not stored on AssemblyAI's database. Once you have finished editing the lambda function, click **Deploy** to save your changes. On the same page, navigate to the **Configuration** section, under **General configuration** adjust the timeout to 15min 0sec and click **Save**. The processing times for transcription will be a lot shorter, but this ensures plenty of time for the function to complete. Lambda General configuration with the timeout set to 15 minutes Now from this page, on the left side panel click **Environment variables**. Click edit and then add an environment variable, `ASSEMBLYAI_API_KEY`, and set the value to your AssemblyAI API key. Then click **Save**. Lambda Environment variables page adding the ASSEMBLYAI_API_KEY variable Now, navigate to the **IAM** services page. On the left side panel under Access Management click **Roles** and search for your Lambda function role (it's structure should look like `function_name-role-id`). Click into the role and then in the **Permissions policies** section click the dropdown for **Add permissions** and then select **Attach policies**. From this page, find the policy named `AmazonS3FullAccess` and click **Add permissions**. IAM role permissions page attaching the AmazonS3FullAccess policy Now, navigate to the **S3** services page and click into the general purpose bucket where your Amazon Connect recordings are stored. Browse to the **Properties** tab and then scroll down to **Event notifications**. Click **Create event notification**. Give the event a name and then in the prefix section, insert the folder path we noted from Step 1 to ensure the event is triggered for the correct folder. S3 Create event notification form with the Call Recordings folder prefix Then in the **Event types** section, select **All object create events**. S3 event notification Event types with All object create events selected Then scroll down to the **Destination** section, set the destination as **Lambda function** and then select the Lambda function we created in Step 2. Then click **Save changes**. S3 event notification Destination set to the Lambda function To finalise the integration, we'll need to set the recording behaviour from within your AWS Contact Flows. Navigate to your Amazon Connect instance access URL and sign in to your Admin account. In the left side panel, navigate to the **Routing** section and then select **Flows**. Choose a flow to test with, in this case we'll utilize the `Sample inbound flow (first contact experience)`. You should see the **Block Library** on the left hand side of the page. In this section, search for `Set recording and analytics behaviour` and then drag the block into your flow diagram and connect the arrows. You can see in our example, we place the block right at the entry of the call flow: Amazon Connect flow with the Set recording and analytics behavior block at the call entry After connecting this block, click the 3 vertical dots in the top right of the block and select **Edit settings**. Scroll down to the **Enable recording and analytics** subsection and expand the **Voice** section. Then select `On` and select `Agent and customer` (or whoever you'd like to record). Then click **Save**, click **Save** again in the top right and then click **Publish** to publish the flow. Set recording block settings enabling Voice recording for Agent and customer With this new flow published, you should now receive recordings for your Amazon Connect calls that utilize that flow, and you should now receive AssemblyAI transcripts for those recordings! The Amazon Connect Call Recordings are saved in the S3 bucket with this naming convention: **/connect/\{instance-name\}/CallRecordings/\{YYYY\}/\{MM\}/\{DD\}/\{contact-id\}\_\{YYYYMMDDThh:mm\}\_UTC.wav** The AssemblyAI Transcripts will be saved in the S3 bucket with this naming convention: **/connect/\{instance-name\}/AssemblyAITranscripts/\{YYYY\}/\{MM\}/\{DD\}/\{contact-id\}\_\{YYYYMMDDThh:mm\}\_UTC.json** To view the logs for this integration, navigate to the **CloudWatch** services page and under the **Logs** section, select **Log groups**. Select the log group that matches your Lambda to view the most recent log stream. --- # Transcribe Genesys Cloud Recordings with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/genesys-cloud Source: docs/integrations/genesys-cloud.mdx Navigation: Overview > Use cases & integrations > Integrations > Telephony tools Description: Transcribe Genesys Cloud Recordings with AssemblyAI documentation. This guide walks through the process of setting up a transcription pipeline to send audio data from Genesys Cloud to AssemblyAI. To accomplish this, we'll stream audio through Genesys's [AudioHook Monitor](https://appfoundry.genesys.com/filter/genesyscloud/listing/a3ff6a99-d866-4734-ab7a-16cff2e4308c) integration to a WebSocket server. Upon call completion, the server will process this audio into a wav file and send it to AssemblyAI's Speech-to-text API for [pre-recorded audio](/pre-recorded-audio) transcription. You can find all the necessary code for this guide in the [example GitHub repository](https://github.com/gsharp-aai/genesys-async-guide). ## Architecture Overview Here's the general flow our app will follow: ``` +---------------------+ +--------------------+ +----------------------+ | 1. Genesys Cloud | → | 2. WebSocket | → | 3. Convert Raw | | (AudioHook Monitor) | | Server | | Audio to WAV | +---------------------+ +--------------------+ +----------------------+ ↓ +--------------------+ +---------------------+ +----------------------+ | 6. S3 Bucket | ← | 5. AssemblyAI API | ← | 4. Audio Upload (S3) | | (Transcript Store) | | (Transcription) | | (Trigger Lambda) | +--------------------+ +---------------------+ +----------------------+ ``` ## Getting started Before we begin, make sure you have: - An AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your [dashboard](https://www.assemblyai.com/dashboard/home). - An AWS account, an [Access key](https://us-east-1.console.aws.amazon.com/iam/home#/security_credentials), and permissions to S3, Lambda, and CloudWatch. - A [Genesys Cloud](https://www.genesys.com/genesys-cloud) account with the necessary permissions to create call flows, phone numbers, and routes. - [ngrok](https://ngrok.com/downloads/mac-os) installed. ## Genesys AudioHook Monitor In order to stream your voice calls to third party services outside of the Genesys Cloud platform, Genesys offers an official integration called [AudioHook Monitor](https://appfoundry.genesys.com/filter/genesyscloud/listing/a3ff6a99-d866-4734-ab7a-16cff2e4308c). This integration allows you to specify the URL of a [WebSocket server](https://developer.mozilla.org/en-US/docs/Web/API/WebSockets_API) that implements the [AudioHook Protocol](https://developer.genesys.cloud/devapps/audiohook/), and once a connection has been established, Genesys will send both text (metadata messages/events encoded as JSON) and binary data (WebSocket frames containing the raw audio data in μ-law (Mu-law, PCMU) format). An understanding of this integration and protocol is recommended before proceeding with this tutorial. Here are some helpful resources to get started: - [Genesys App Foundry](https://appfoundry.genesys.com/filter/genesyscloud/listing/a3ff6a99-d866-4734-ab7a-16cff2e4308c) - [AudioHook Monitor](https://help.mypurecloud.com/articles/about-audiohook-monitor/) - [AudioHook Protocol](https://developer.genesys.cloud/devapps/audiohook/) - [AudioHook Sample Service Repo](https://github.com/purecloudlabs/audiohook-reference-implementation#genesys-audiohook-sample-service) ## Step 1: Create a call flow in Genesys (optional) You may already have an inbound call flow set up in Genesys (if so, skip to [Step 3](/integrations/genesys-cloud#step-3-create-a-s3-bucket)), but we'll create a simple one from scratch for the sake of this tutorial. **Adapting existing call flows** All you need to add is the **Audio Monitoring** step from the toolbox and make sure **Suppress recording for the entire flow** is unchecked in the flow's **Recording and Speech Recognition** settings. Within the [Architect tool](https://apps.usw2.pure.cloud/architect/#/inboundcall/flows), click **Add** to create a new call flow. Enter a **Name** for your flow and click **Create Flow**. Select the newly created flow to open the drag and drop editor. Create Flow dialog for adding a new inbound call flow in Genesys Architect Create a **Reusable Task** from the bottom left of the left-side menu. From the Toolbox, search for **Audio Monitoring** and drag it just after the **Start** step of our flow. In the right-side menu for this option, make sure **Enable Monitoring** is enabled. Audio Monitoring step added after Start with Enable Monitoring turned on in Genesys Architect Back in the Toolbox, search for **Transfer to User** and set that as the next step. In the right-hand menu, under **User** select a caller. Under **Pre-Transfer Audio** and **Failed Transfer Audio**, type your preferred messages. Transfer to User step configured with pre-transfer and failed transfer audio in Genesys Architect Search for **Disconnect** in the toolbox and drag that as the step following **Failure**. Disconnect step placed after the Failure path in the Genesys call flow Search the Toolbox for **Jump to Reusable Task** and drag this tool to the Main Menu at the top of the left-side menu. Select a **DTMF** and **Speech Recognition** value (this will be used to transfer your call to the agent). Under **Task**, select the task you just created. Jump to Reusable Task tool added to the Main Menu with DTMF and speech recognition values Under Settings in the left-side menu, navigate to the **Recording and Speech Recognition** section. Make sure **Suppress recording for the entire flow** is unchecked. Recording and Speech Recognition settings with Suppress recording for the entire flow unchecked In the top navbar, make sure to click **Save** and then click **Publish** to have your changes take effect. ## Step 2: Setup a phone and routing for your flow (optional) In the Genesys Cloud Admin section, navigate to the [Phones page](https://apps.usw2.pure.cloud/directory/#/admin/telephony/phone-management/phones) under the **Telephony section** and click **Add** to create a new phone. For **Person**, assign the User from your organization that you selected for the **Transfer to User** step in the [previous section](/integrations/genesys-cloud#step-1-create-a-call-flow-in-genesys-optional). Creating a new phone on the Genesys Cloud Phone Management page Assigning the transfer-to user as the Person for the new phone in Genesys Cloud Navigate to the [Number Management](https://apps.usw2.pure.cloud/directory/#/admin/telecom/numbers/numbers) page under the **Genesys Cloud Voice** section and select **Purchase Numbers**. Enter an area code and click **Search**. Select a phone number from the list and click **Complete Purchase**. Number Management page with the Purchase Numbers option in Genesys Cloud Voice Searching for a phone number by area code before completing purchase in Genesys Cloud Navigate to the [Call Routing](https://apps.usw2.pure.cloud/directory/#/admin/routing/ivrs) page under the **Routing** section and select **Add**. Under **What call flow should be used?** select your flow. For **Inbound Numbers**, type the number you purchased in the above step. Then click **Create**. Call Routing page with the Add option under the Routing section in Genesys Cloud Creating a call route mapping the purchased inbound number to the call flow Under the **Telephony** section, navigate to the [External Trunks](https://apps.usw2.pure.cloud/directory/#/admin/telephony/trunks/external) page. Click **Create New**. Under **Caller ID**, the **Caller Address** will be an E.164 number and the phone number you created. Creating a new external trunk on the Genesys Cloud External Trunks page Setting the trunk Caller ID caller address to the purchased E.164 phone number Under **SIP Access Control**, select **Allow All** (_note: this is only for development and testing purposes, please specify actual IPs in production_). Under the **Media** section, make sure you select **Record calls on this trunk**. Then click **Save External Trunk** (it may take a few moments for your trunk to be ready). SIP Access Control set to Allow All in the external trunk settings Media section with Record calls on this trunk enabled for the external trunk ## Step 3: Create a S3 bucket After our Genesys call ends, store the audio file in S3. Click **Create bucket**. Give your bucket a name like `your-audiohook-bucket`. Scroll down and click **Create bucket**. Creating an S3 bucket named your-audiohook-bucket in the AWS console ## Step 4: Create a WebSocket server In this step, we'll set up a WebSocket server to receive messages and audio data from Genesys as they are sent. Our server must respond to certain events (i.e. `open`, `close`, `ping`, `pause`, etc.) according to the AudioHook protocol. Outside of these event messages, audio data is also transferred. We'll capture and temporarily store this audio locally until the connection is closed, at which point the server processes the audio to a `wav` file and uploads both the `wav` and `raw` audio files to a S3 bucket. The AudioHook Monitor will send requests to a WebSocket URL we specify when setting up the integration in [Step 5](/integrations/genesys-cloud#step-5-setting-up-audiohook-monitor). When first enabled, AudioHook Monitor will do a quick verification step to ensure that the WebSocket server has implemented the AudioHook protocol correctly. For this example, the server is written in JavaScript ([Express](https://expressjs.com/)) and hosted locally. We'll use [ngrok](https://ngrok.com/) to create a secure tunnel that exposes it to the internet with a public URL so that Genesys can make a connection. However, the server can be implemented using your preferred programming language and deployed in whatever environment you choose, provided both support WebSocket TLS connections for secure bidirectional text and binary message exchange. **Server implementation** This server is a method to get up and running quickly for development and testing purposes without the complexity of a production deployment. How you implement this in practice will vary widely depending on your traffic volume, scaling needs, reliability requirements, security concerns, etc. Clone this [example repo](https://github.com/gsharp-aai/genesys-async-guide) of a WebSocket server that implements the AudioHook protocol. Follow the `README` instructions to download the necessary dependencies and start the server. Make sure to look over `server.js` to get an understanding of how the requests from Genesys are received, processed, and responded to, as well as how the audio is stored, converted, and uploaded to our S3 bucket. Make sure to create a `.env` file and set the variables: ```env PORT=3000 AWS_REGION=us-east-1 # Region of your S3 bucket AWS_ACCESS_KEY_ID=YOUR_ACCESS_KEY_ID # Found under IAM > Security Credentials AWS_SECRET_ACCESS_KEY=YOUR_SECRET_ACCESS_KEY # Found under IAM > Security Credentials S3_BUCKET=your-audiohook-bucket # Name of your S3 bucket S3_KEY_PREFIX=calls/ # The file structure you want your bucket to follow API_KEY=YOUR_API_KEY # Used for authenticating messages from Genesys RECORDINGS_DIR=./recordings # Temp file storage location ``` `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` can be found on your account's [IAM > Security Credentials page](https://us-east-1.console.aws.amazon.com/iam/home?region=us-east-1#/security_credentials). `API_KEY` is explained further in the [next step](/integrations/genesys-cloud#step-5-setting-up-audiohook-monitor), but it can be anything you want to verify that the requests are actually originating from Genesys. Download [ngrok](https://ngrok.com/). Assuming your server is running on `port 3000`, run `ngrok http 3000 --inspect=false` in your terminal. From the resulting terminal output, note the forwarding url that should look something like this: `https://.ngrok-free.app`. This is our WebSocket server URL that we'll provide to the AudioHook Monitor in the next step. ngrok terminal output showing the forwarding URL for the local WebSocket server ## Step 5: Setting up AudioHook Monitor In the Genesys Cloud Admin section, navigate to the [Integrations](https://apps.usw2.pure.cloud/directory/#/admin/integrations/apps) page. Add a new integration via the plus sign in the top right corner. Adding a new integration from the Genesys Cloud Integrations page Search for **AudioHook Monitor** and install. Installing the AudioHook Monitor integration from the Genesys Cloud integration catalog Navigate to the AudioHook Monitor's **Configuration** tab. Under the **Properties section**, make sure both channels are selected and the **Connection URI** is set to the ngrok url from the [previous step](/integrations/genesys-cloud#step-4-create-a-websocket-server). For the ngrok url, replace `https` with `wss`. AudioHook Monitor Configuration tab with the wss Connection URI and both channels selected In the **Configuration** tab, navigate to the **Credentials** section, and click **Configure**. Here you can set an API key to a value our server will use to verify that requests originated from Genesys. This is done via the `X-API-KEY` request header. Our server will compare this key to the value we set in our `.env` for `API_KEY`, so make sure they match. Click **Save**. Setting the X-API-KEY credential in the AudioHook Monitor Credentials configuration Back on the **Integrations** page, click the toggle button under the **Status** column to activate your AudioHook. Genesys will attempt to verify our server is correctly configured according to the AudioHook protocol. If it is unable to do so, a red error will show with the reason for the failed connection. If it succeeds, the connection will toggle to Active. AudioHook Monitor integration toggled to Active status after successful verification ## Step 6: Set up your AssemblyAI API call Navigate to the [Lambda](https://us-east-1.console.aws.amazon.com/lambda/home) services page, and create a new function. Set the runtime to `Node.js 22.x`. In the **Change default execution role** section, choose the option to create a **new role with basic Lambda permissions**. Assign a function name and then click **Create function**. Creating a new Node.js Lambda function with a basic execution role in the AWS console In this new function, scroll down to the **Code Source** section and paste the following code into `index.js`: ```javascript expandable // Import required AWS SDK modules import { S3 } from "@aws-sdk/client-s3"; import { getSignedUrl } from "@aws-sdk/s3-request-presigner"; import { GetObjectCommand } from "@aws-sdk/client-s3"; // Configure logging const logger = { info: (data) => console.log(JSON.stringify(data)), error: (data) => console.error(JSON.stringify(data)), }; // Configuration settings for AssemblyAI // See config parameters here: https://www.assemblyai.com/docs/api-reference/transcripts/submit const ASSEMBLYAI_CONFIG = { multichannel: true, // Using multichannel here as we told Genesys to send us multichannel audio. }; // Initialize AWS S3 client const s3Client = new S3(); /** * Generate a presigned URL for the S3 object * @param {string} bucket - S3 bucket name * @param {string} key - S3 object key * @param {number} expiration - URL expiration time in seconds * @returns {Promise} Presigned URL */ const getPresignedUrl = async (bucket, key, expiration = 3600) => { logger.info({ message: "Generating presigned URL", bucket: bucket, key: key, expiration: expiration, }); const command = new GetObjectCommand({ Bucket: bucket, Key: key, }); return getSignedUrl(s3Client, command, { expiresIn: expiration }); }; /** * Delete transcript data from AssemblyAI's database * @param {string} transcriptId - The AssemblyAI transcript ID to delete * @param {string} apiKey - The AssemblyAI API key * @returns {Promise} True if deletion was successful, False otherwise */ const deleteTranscriptFromAssemblyAI = async (transcriptId, apiKey) => { try { const response = await fetch( `https://api.assemblyai.com/v2/transcript/${transcriptId}`, { method: "DELETE", headers: { authorization: apiKey, "content-type": "application/json", }, } ); if (response.ok) { logger.info( `Successfully deleted transcript ${transcriptId} from AssemblyAI` ); return true; } else { const errorData = await response.text(); logger.error( `Failed to delete transcript ${transcriptId}: HTTP ${response.status} - ${errorData}` ); return false; } } catch (error) { logger.error(`Error deleting transcript ${transcriptId}: ${error.message}`); return false; } }; /** * Submit audio for transcription * @param {object} requestData - Request data including audio URL and config * @param {string} apiKey - AssemblyAI API key * @returns {Promise} Transcript ID */ const submitTranscriptionRequest = async (requestData, apiKey) => { const response = await fetch("https://api.assemblyai.com/v2/transcript", { method: "POST", headers: { authorization: apiKey, "content-type": "application/json", }, body: JSON.stringify(requestData), }); if (!response.ok) { const errorText = await response.text(); throw new Error(`Failed to submit audio for transcription: ${errorText}`); } const responseData = await response.json(); const transcriptId = responseData.id; logger.info({ message: "Audio submitted for transcription", transcript_id: transcriptId, }); return transcriptId; }; /** * Poll for transcription completion * @param {string} transcriptId - Transcript ID * @param {string} apiKey - AssemblyAI API key * @returns {Promise} Transcription data */ const pollTranscriptionStatus = async (transcriptId, apiKey) => { const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms)); // Keep polling until we get a completion or error while (true) { const response = await fetch( `https://api.assemblyai.com/v2/transcript/${transcriptId}`, { method: "GET", headers: { authorization: apiKey, "content-type": "application/json", }, } ); if (!response.ok) { const errorText = await response.text(); throw new Error(`Failed to poll transcription status: ${errorText}`); } const pollingData = await response.json(); if (pollingData.status === "completed") { logger.info({ message: "Transcription completed successfully" }); return pollingData; } else if (pollingData.status === "error") { throw new Error(`Transcription failed: ${pollingData.error}`); } // Wait before polling again await sleep(3000); } }; /** * Transcribe audio using AssemblyAI API * @param {string} audioUrl - URL of the audio file * @param {string} apiKey - AssemblyAI API key * @returns {Promise} Transcription data */ const transcribeAudio = async (audioUrl, apiKey) => { logger.info({ message: "Starting audio transcription" }); // Prepare request data with config parameters const requestData = { audio_url: audioUrl, ...ASSEMBLYAI_CONFIG }; // Submit the audio file for transcription const transcriptId = await submitTranscriptionRequest(requestData, apiKey); // Poll for transcription completion return await pollTranscriptionStatus(transcriptId, apiKey); }; /** * Lambda function handler * @param {object} event - S3 event * @param {object} context - Lambda context * @returns {Promise} Response */ export const handler = async (event, context) => { try { // Get the AssemblyAI API key from environment variables const apiKey = process.env.ASSEMBLYAI_API_KEY; if (!apiKey) { throw new Error("ASSEMBLYAI_API_KEY environment variable is not set"); } // Process each record in the S3 event const records = event.Records || []; for (const record of records) { // Get the S3 bucket and key const bucket = record.s3.bucket.name; const key = decodeURIComponent(record.s3.object.key.replace(/\+/g, " ")); // Generate a presigned URL for the audio file const audioUrl = await getPresignedUrl(bucket, key); // Get the full transcript JSON from AssemblyAI const transcriptData = await transcribeAudio(audioUrl, apiKey); // Prepare the transcript key - maintaining path structure but changing directory and extension const transcriptKey = key .replace("audio", "transcripts", 1) .replace(".wav", ".json"); // Convert the JSON data to a string const transcriptJsonStr = JSON.stringify(transcriptData, null, 2); // Upload the transcript JSON to the same bucket but in transcripts directory await s3Client.putObject({ Bucket: bucket, // Use the same bucket Key: transcriptKey, // Store under the /transcripts directory Body: transcriptJsonStr, ContentType: "application/json", }); logger.info({ message: "Transcript uploaded to transcript bucket successfully.", key: transcriptKey, }); // Uncomment the following line to delete transcript data from AssemblyAI after saving to S3 // https://www.assemblyai.com/docs/api-reference/transcripts/delete // await deleteTranscriptFromAssemblyAI(transcriptData.id, apiKey); } return { statusCode: 200, body: JSON.stringify({ message: "Audio file(s) processed successfully", detail: "Transcripts have been stored in the AssemblyAITranscripts directory", }), }; } catch (error) { console.error(`Error: ${error.message}`); return { statusCode: 500, body: JSON.stringify({ message: "Error processing audio file(s)", error: error.message, }), }; } }; ``` At the top of the Lambda function, you can edit the config to enable features for your transcripts. Since our call is two channels, we'll want to set `multichannel` to `true`. To see all available parameters, check out our [API reference](/api-reference/transcripts/submit). ```javascript ASSEMBLYAI_CONFIG = { multichannel: true, // 'language_code': 'en_us', // 'redact_pii': true // etc. }; ``` If you would like to delete transcripts from AssemblyAI after completion, you can uncomment `line 212` to enable the `deleteTranscriptFromAssemblyAI` function. This ensures the transcript data is only saved to your S3 bucket and not stored on AssemblyAI's database. Once you have finished editing the Lambda function, click **Deploy** to save your changes. Lambda code source editor with the transcription handler pasted into index.js On the same page, navigate to the **Configuration** section. Under **General configuration**, click **Edit**, and then adjust **Timeout** to `15min 0sec` and click **Save**. The processing times for transcription will be much shorter, but this ensures the function will have plenty of time to run. Lambda Configuration tab General configuration section in the AWS console Setting the Lambda function timeout to 15 minutes in general configuration On the left side panel, click **Environment variables**. Click **Edit**. Add an environment variable, `ASSEMBLYAI_API_KEY`, and set the value to your AssemblyAI [API key](https://www.assemblyai.com/dashboard/home). Then click **Save**. Adding the ASSEMBLYAI_API_KEY environment variable to the Lambda function Now, navigate to the [IAM](https://us-east-1.console.aws.amazon.com/iam/) services page. On the left side panel under **Access Management**, click **Roles** and search for your Lambda function's role (its structure should look like `-`). Click the role and then in the **Permissions policies** section click the dropdown for **Add permissions** and then select **Attach policies.** Attaching policies to the Lambda function's IAM role in the AWS console From this page, find the policies named `AmazonS3FullAccess` and `CloudWatchEventsFullAccess`. Click **Add permissions** for both. Adding AmazonS3FullAccess and CloudWatchEventsFullAccess permissions to the IAM role `CloudWatchEventsFullAccess` is optional, but helpful for debugging purposes. Once your Lambda runs, it should output all logs to [CloudWatch](https://us-east-1.console.aws.amazon.com/cloudwatch) under a Log group `/aws/lambda/` Now, navigate to the [S3](https://us-east-1.console.aws.amazon.com/s3) services page and click into the general purpose bucket where your Genesys recordings are stored. Browse to the **Properties** tab and then scroll down to **Event notifications**. Click **Create event notification**. Creating an S3 event notification from the bucket Properties tab in the AWS console Give the event a name and then in the **Prefix** section enter `calls/` (or whatever `S3_KEY_PREFIX` is set to), and in the **Suffix** section enter `.wav`. This will ensure the event is triggered once our `wav` file has been uploaded. In the **Event types** section, select **All object create events**. Configuring the S3 event notification with calls/ prefix, .wav suffix, and object create events Scroll down to the **Destination** section, set the destination as **Lambda function** and then select the Lambda function we created in [Step 6](/integrations/genesys-cloud#step-6-set-up-your-assemblyai-api-call). Then click **Save changes**. Setting the S3 event notification destination to the Lambda function ## Step 7: Transcribe your first call To test everything is working, call the phone number you linked to this flow in [Step 2](/integrations/genesys-cloud#step-2-setup-a-phone-and-routing-for-your-flow-optional). Referring to the example flow above, press the DTMF value on the key pad or say the Speech Recognition value. Once transferred, your WebSocket server should start to receive data and output to console: ```bash { version: '2', id: '', type: 'ping', seq: 4, position: 'PT8.2S', parameters: { rtt: 'PT0.035392266S' }, serverseq: 3 } Received binary audio data: 3200 bytes Received binary audio data: 3200 bytes ... Processed 146KB of audio data so far ``` Once the call has ended, you should see the following server logs: ```bash expandable { version: '2', id: '', type: 'close', seq: 5, position: 'PT10.2S', parameters: { reason: 'end' }, serverseq: 4 } Handling close message Closing file stream Converting raw audio to WAV: '' # Skipping ffmpeg output for brevity... Uploading recording '' to S3 Successfully uploaded raw recording to S3: '' Successfully uploaded WAV recording to S3: '' Sent closed response, seq=5 WebSocket closed for session '': code=1000, reason=Session Ended Cleaning up session '' Deleted local raw recording file: '' Deleted local WAV recording file: '' ``` To view the logs for this Lambda function, navigate to the [CloudWatch](https://us-east-1.console.aws.amazon.com/cloudwatch) services page and under the Logs section, select **Log groups**. Select the log group that matches your Lambda to view the most recent log stream. This can be very useful for debugging purposes if you run into any issues. CloudWatch Log groups page for viewing the Lambda function's log streams Head to your S3 bucket. Within the `/calls` directory, files will be stored under a unique identifier with the following structure: `your-audiohook-bucket/calls/__/` with audio files (both `raw` and `wav`) under `/audio` and transcript responses under `/transcripts`. S3 bucket showing stored call recordings and transcript files after a successful run The `raw` file can be nice to have for conversions to other formats in the future, but this step can be omitted to save on storage costs. **Success!** You have successfully integrated AssemblyAI with Genesys Cloud via AudioHook Monitor. If you run into any issues or have further questions, please reach out to our [Support team](https://www.assemblyai.com/contact/support). ## Other considerations ### Supported audio formats - Audio is sent as binary WebSocket frames containing the raw audio data in the negotiated format. Currently, only μ-law (Mu-law, PCMU) is [supported](https://developer.genesys.cloud/devapps/audiohook/session-walkthrough#audio-streaming). - Before being uploaded to S3, the audio is converted to `wav` format using [ffmpeg](https://ffmpeg.org/). As a lossless format, `wav` generally results in high transcription accuracy, but is not required. A full list of [file formats supported by AssemblyAI's API](/faq/what-audio-and-video-file-types-are-supported-by-your-api) can be found in our FAQ. ### Multichannel - As mentioned in [Step 6](/integrations/genesys-cloud#step-6-set-up-your-assemblyai-api-call), the `multichannel` parameter should be enabled as the files are stereo utilizing one channel for each participant. When possible, multiple channels are recommended by AssemblyAI for the most accurate transcription results. - If single channel is preferred, you can simplify the approach to only send a single channel with both speakers (via AudioHook) and adjust your server code to be single channel. --- # \U0001F99C️\U0001F517 LangChain Integration with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/langchain Source: docs/integrations/langchain.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools > LangChain Description: Transcribe audio in LangChain using the built-in integration with AssemblyAI. ## AssemblyAI LangChain Integration AssemblyAI has a LangChain integration for both Python and JavaScript. Learn more about our language-specific integrations: ## New to LangChain? [LangChain](https://www.langchain.com) is an open-source framework for developing applications with [Large Language Models (LLMs)](https://www.assemblyai.com/blog/introduction-large-language-models-generative-ai/) and other AI technologies. LangChain has a set of pre-built components that you can use to load data and apply LLMs to your data. However, LLMs only operate on textual data and don't understand speech in audio and video files. To apply LLMs to speech, you first need to transcribe the audio to text, which is what the AssemblyAI integration helps you with. --- # \U0001F99C️\U0001F517 LangChain Python Integration with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/langchain/python Source: docs/integrations/langchain/python.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools > LangChain Description: Transcribe audio in LangChain Python using the built-in integration with AssemblyAI. To apply LLMs to speech, you first need to transcribe the audio to text, which is what the AssemblyAI integration for LangChain helps you with. Looking for the LangChain JavaScript integration?
[Go to the LangChain.JS integration](/integrations/langchain/js). ## Quickstart Install [the AssemblyAI package](https://github.com/langchain-ai/langchain) and [the AssemblyAI Python SDK](https://github.com/AssemblyAI/assemblyai-python-sdk): ```bash pip install langchain pip install assemblyai ``` Set your AssemblyAI API key as an environment variable named `ASSEMBLYAI_API_KEY`. You can [get a free AssemblyAI API key from the AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home). ```bash # Mac/Linux: export ASSEMBLYAI_API_KEY=YOUR_API_KEY # Windows: set ASSEMBLYAI_API_KEY=YOUR_API_KEY ``` Import the `AssemblyAIAudioTranscriptLoader` from `langchain.document_loaders`. ```python from langchain.document_loaders import AssemblyAIAudioTranscriptLoader ``` 1. Pass the local file path or URL as the `file_path` argument of the `AssemblyAIAudioTranscriptLoader`. 2. Call the `load` method to get the transcript as LangChain documents. ```python audio_file = "https://assembly.ai/sports_injuries.mp3" # or a local file path: audio_file = "./sports_injuries.mp3" loader = AssemblyAIAudioTranscriptLoader(file_path=audio_file) docs = loader.load() ``` The `load` method returns an array of documents, but by default, there's only one document in the array with the full transcript. The transcribed text is available in the `page_content` attribute: ```python docs[0].page_content # Load time, a new president and new congressional makeup. Same old ... ``` The `metadata` contains the full JSON response with more meta information: ```python { 'language_code': , 'audio_url': 'https://assembly.ai/nbc.mp3', 'punctuate': True, 'format_text': True, ... } ``` ## Transcript formats You can specify the `transcript_format` argument to load the transcript in different formats. Depending on the format, `load_data()` returns either one or more documents. These are the different `TranscriptFormat` options: - `TEXT`: One document with the transcription text - `SENTENCES`: Multiple documents, splits the transcription by each sentence - `PARAGRAPHS`: Multiple documents, splits the transcription by each paragraph - `SUBTITLES_SRT`: One document with the transcript exported in SRT subtitles format - `SUBTITLES_VTT`: One document with the transcript exported in VTT subtitles format ```python import assemblyai as aai from langchain.document_loaders import AssemblyAIAudioTranscriptLoader from langchain.document_loaders.assemblyai import TranscriptFormat loader = AssemblyAIAudioTranscriptLoader( file_path="./your_file.mp3", transcript_format=TranscriptFormat.SENTENCES, ) docs = loader.load() ``` ## Transcription config You can also specify the `config` argument to use different transcript features and speech understanding models. Here's an example of using the `config` argument to enable speaker labels, auto chapters, and entity detection: ```python import assemblyai as aai from langchain.document_loaders import AssemblyAIAudioTranscriptLoader config = aai.TranscriptionConfig( speaker_labels=True, auto_chapters=True, entity_detection=True ) loader = AssemblyAIAudioTranscriptLoader(file_path="./your_file.mp3", config=config) ``` For the full list of options, see [Transcript API reference](/api-reference/transcripts/submit#request). ## Pass the AssemblyAI API key as an argument Instead of configuring the AssemblyAI API key as the `ASSEMBLYAI_API_KEY` environment variable, you can also pass it as the `api_key` argument. ```python loader = AssemblyAIAudioTranscriptLoader( file_path="./your_file.mp3", api_key="" ) ``` ## Additional resources You can learn more about using LangChain with AssemblyAI in these resources. - [LangChain docs for the AssemblyAI document loader](https://python.langchain.com/docs/integrations/document_loaders/assemblyai) - [How to use audio data in LangChain with Python](https://www.assemblyai.com/blog/load-audio-langchain-python/) - [Retrieval Augmented Generation on audio data with LangChain and Chroma](https://www.assemblyai.com/blog/retrieval-augmented-generation-audio-langchain/) - [Build LangChain Audio Apps with Python in 5 Minutes](https://www.youtube.com/watch?v=7w7ysaDz2W4) - [How to use LangChain for RAG over audio files](https://www.youtube.com/watch?v=l9YJrLg61ac) - [AssemblyAI Python SDK](https://github.com/AssemblyAI/assemblyai-python-sdk) --- # \U0001F99C️\U0001F517 LangChain JavaScript Integration with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/langchain/js Source: docs/integrations/langchain/js.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools > LangChain Description: Transcribe audio in LangChain.JS using the built-in integration with AssemblyAI. To apply LLMs to speech, you first need to transcribe the audio to text, which is what the AssemblyAI integration for LangChain helps you with. Looking for the Python integration?
[Go to the LangChain Python integration](/integrations/langchain/python). ## Quickstart Add the [AssemblyAI SDK](https://github.com/AssemblyAI/assemblyai-node-sdk) to your project: ```bash npm install langchain @langchain/community ``` ```bash yarn add langchain @langchain/community ``` ```bash pnpm add langchain @langchain/community ``` ```bash bun add langchain @langchain/community ``` To use the loaders, you need an [AssemblyAI account](https://www.assemblyai.com/dashboard/signup) and get your AssemblyAI API key from the [dashboard](https://www.assemblyai.com/dashboard/home). Configure the API key as the `ASSEMBLYAI_API_KEY` environment variable or the `apiKey` options parameter. ```javascript expandable import { AudioTranscriptLoader, // AudioTranscriptParagraphsLoader, // AudioTranscriptSentencesLoader } from "@langchain/community/document_loaders/web/assemblyai"; // You can also use a local file path and the loader will upload it to AssemblyAI for you. const audioUrl = "https://assembly.ai/espn.m4a"; // Use `AudioTranscriptParagraphsLoader` or `AudioTranscriptSentencesLoader` for splitting the transcript into paragraphs or sentences const loader = new AudioTranscriptLoader( { audio: audioUrl, // any other parameters as documented here: https://www.assemblyai.com/docs/api-reference/transcript#create-a-transcript }, { apiKey: "", // or set the `ASSEMBLYAI_API_KEY` env variable } ); const docs = await loader.load(); console.dir(docs, { depth: Infinity }); ``` - You can use the `AudioTranscriptParagraphsLoader` or `AudioTranscriptSentencesLoader` to split the transcript into paragraphs or sentences. - If the `audio_file` is a local file path, the loader will upload it to AssemblyAI for you. - The `audio_file` can also be a video file. See the [list of supported file types in the FAQ doc](/faq/what-audio-and-video-file-types-are-supported-by-your-api). - If you don't pass in the `apiKey` option, the loader will use the `ASSEMBLYAI_API_KEY` environment variable. - You can add more properties in addition to `audio`. Find the full list of request parameters in the [AssemblyAI API docs](/api-reference/overview).
You can also use the `AudioSubtitleLoader` to get `srt` or `vtt` subtitles as a document. ```javascript import { AudioSubtitleLoader } from "@langchain/community/document_loaders/web/assemblyai"; // You can also use a local file path and the loader will upload it to AssemblyAI for you. const audioUrl = "https://assembly.ai/espn.m4a"; const loader = new AudioSubtitleLoader( { audio: audioUrl, // any other parameters as documented here: https://www.assemblyai.com/docs/api-reference/transcript#create-a-transcript }, "srt", // srt or vtt { apiKey: "", // or set the `ASSEMBLYAI_API_KEY` env variable } ); const docs = await loader.load(); console.dir(docs, { depth: Infinity }); ``` ## Additional resources You can learn more about using LangChain with AssemblyAI in these resources: - [The LangChain docs for the AssemblyAI document loader](https://js.langchain.com/docs/integrations/document_loaders/web_loaders/assemblyai_audio_transcription) - [How to integrate spoken audio into LangChain.js using AssemblyAI](https://www.assemblyai.com/blog/integrate-audio-langchainjs/) - [Integrate Audio into LangChain.js apps in 5 Minutes](https://www.youtube.com/watch?v=hNpUSaYZIzs) - [AssemblyAI JavaScript SDK](https://github.com/AssemblyAI/assemblyai-node-sdk) --- # ▲ Vercel AI SDK Integration with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/vercel-ai-sdk Source: docs/integrations/vercel-ai-sdk.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Transcribe audio with the Vercel AI SDK using the official @ai-sdk/assemblyai provider. The [AI SDK](https://ai-sdk.dev) gives you a single, unified `transcribe()` API that works across speech-to-text providers. The [`@ai-sdk/assemblyai`](https://ai-sdk.dev/providers/ai-sdk-providers/assemblyai) provider — maintained by Vercel in the [`vercel/ai`](https://github.com/vercel/ai) repo — plugs AssemblyAI's speech models into that API, so you can transcribe audio in any TypeScript or JavaScript project without calling the AssemblyAI API directly. ## Quickstart Install the AI SDK core package along with the AssemblyAI provider: ```bash npm install ai @ai-sdk/assemblyai ``` ```bash yarn add ai @ai-sdk/assemblyai ``` ```bash pnpm add ai @ai-sdk/assemblyai ``` ```bash bun add ai @ai-sdk/assemblyai ``` You'll need an [AssemblyAI account](https://www.assemblyai.com/dashboard/signup) and an API key from your [dashboard](https://www.assemblyai.com/dashboard/home). The provider reads it from the `ASSEMBLYAI_API_KEY` environment variable: ```bash export ASSEMBLYAI_API_KEY="" ``` Then transcribe a local file with `transcribe()`: ```typescript import { transcribe } from 'ai'; import { assemblyai } from '@ai-sdk/assemblyai'; import { readFile } from 'node:fs/promises'; const result = await transcribe({ model: assemblyai.transcription('universal-3-5-pro'), audio: await readFile('audio.mp3'), }); console.log(result.text); ``` `transcribe` is a stable export as of AI SDK v7, which is what `npm install ai` installs today. The examples on this page target v7. **Using AI SDK v6?** The examples on this page require v7. AI SDK v6 (and v5) export only `experimental_transcribe` — `import { transcribe } from 'ai'` throws on those versions — and the latest provider (`@ai-sdk/assemblyai` 3.x) is built for v7, so it won't work against v6. If your project is on v6 (for example, an existing project or a registry policy that hasn't picked up v7 yet), install the v6-compatible packages and alias the experimental export: ```bash npm install ai@ai-v6 @ai-sdk/assemblyai@ai-v6 ``` ```typescript import { experimental_transcribe as transcribe } from 'ai'; ``` Everything else on this page — the `providerOptions`, model IDs, and result shape — is the same. ## Configuring the provider The default `assemblyai` instance reads your key from `ASSEMBLYAI_API_KEY`. To pass the key explicitly — for example, from a secret manager or in a multi-tenant setup — or to add custom headers or a custom `fetch` implementation, create your own provider instance with `createAssemblyAI`: ```typescript import { createAssemblyAI } from '@ai-sdk/assemblyai'; const assemblyai = createAssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY, }); const result = await transcribe({ model: assemblyai.transcription('universal-3-5-pro'), audio: await readFile('audio.mp3'), }); ``` ## Choosing a speech model `universal-3-5-pro` is the recommended default. `universal-2` is also supported. Pass the model ID to `assemblyai.transcription()`: ```typescript const model = assemblyai.transcription('universal-3-5-pro'); ``` For a breakdown of each model's accuracy, latency, and language support, see [Select the speech model](/pre-recorded-audio/select-the-speech-model). ## Using AssemblyAI features AssemblyAI's diarization and audio-intelligence features are enabled through `providerOptions.assemblyai`. Option keys are the **camelCase** equivalents of the API's snake_case parameters (for example, `speaker_labels` → `speakerLabels`, `redact_pii` → `redactPii`): ```typescript import { transcribe } from 'ai'; import { assemblyai, type AssemblyAITranscriptionModelOptions, } from '@ai-sdk/assemblyai'; import { readFile } from 'node:fs/promises'; const result = await transcribe({ model: assemblyai.transcription('universal-3-5-pro'), audio: await readFile('meeting.mp3'), providerOptions: { assemblyai: { speakerLabels: true, keytermsPrompt: ['AssemblyAI', 'Universal'], redactPii: true, redactPiiPolicies: ['person_name', 'email_address'], } satisfies AssemblyAITranscriptionModelOptions, }, }); ``` Annotating the options object with `satisfies AssemblyAITranscriptionModelOptions` gives you editor autocomplete and type-checking for the full option set, so you don't have to memorize parameter names or hunt through docs. The table below reflects `@ai-sdk/assemblyai` 3.0.5. For the always-current list — including any options added in newer releases — see the [`@ai-sdk/assemblyai` provider reference](https://ai-sdk.dev/providers/ai-sdk-providers/assemblyai), which lives alongside the provider code and is the canonical source. ### Transcription options Every option below is optional and passed under `providerOptions.assemblyai`. Types and behavior follow the AssemblyAI transcription API. | Option | Type | Description | | --- | --- | --- | | `audioEndAt` | number | End time of the audio, in milliseconds. | | `audioStartFrom` | number | Start time of the audio, in milliseconds. | | `autoChapters` | boolean | Automatically generate chapters for the transcript. | | `autoHighlights` | boolean | Automatically generate highlights for the transcript. | | `boostParam` | enum | Boost level for `wordBoost`. Allowed values: `low`, `default`, `high`. **Deprecated** — applies only to the deprecated `wordBoost`; use `keytermsPrompt` instead. | | `contentSafety` | boolean | Enable content safety filtering. | | `contentSafetyConfidence` | number | Confidence threshold for content safety filtering (25-100). | | `customSpelling` | array of objects | Custom spelling rules. Each object has `from` (array of strings) and `to` (string). | | `disfluencies` | boolean | Include disfluencies (um, uh, etc.) in the transcript. | | `domain` | string | Enable a domain-specific model for specialized terminology. Currently supports `medical-v1` (Medical Mode). | | `entityDetection` | boolean | Detect entities in the transcript. | | `filterProfanity` | boolean | Filter profanity in the transcript. | | `formatText` | boolean | Apply text formatting to the transcript. | | `iabCategories` | boolean | Include IAB categories in the transcript. | | `keytermsPrompt` | array of strings | Domain-specific keyterms to boost recognition for (max 6 words per phrase). Replaces `wordBoost` for newer models — supported by `universal-3-pro`, `universal-3-5-pro`, and `slam-1` (and `universal-2` when enabled). | | `languageCode` | string | Language code for the audio. Supports numerous ISO-639-1 and ISO-639-3 codes. | | `languageConfidenceThreshold` | number | Confidence threshold for language detection. | | `languageDetection` | boolean | Enable automatic language detection. | | `languageDetectionOptions` | object | Options for language detection: `expectedLanguages` (array of strings), `fallbackLanguage` (string), `codeSwitching` (boolean), `codeSwitchingConfidenceThreshold` (number, 0-1). | | `multichannel` | boolean | Process multiple audio channels separately. | | `prompt` | string | Natural-language context (up to 1,500 words) to steer the model. Only supported by `universal-3-pro`, `universal-3-5-pro`, and `slam-1`. | | `punctuate` | boolean | Add punctuation to the transcript. | | `redactPii` | boolean | Redact personally identifiable information (PII). | | `redactPiiAudio` | boolean | Redact PII in the audio file. | | `redactPiiAudioOptions` | object | Options for PII-redacted audio: `returnRedactedNoSpeechAudio` (boolean), `overrideAudioRedactionMethod` (`silence`). Requires `redactPiiAudio`. | | `redactPiiAudioQuality` | enum | Quality of the redacted audio file. Allowed values: `mp3`, `wav`. | | `redactPiiPolicies` | array of enums | Which types of information to redact (e.g. `person_name`, `phone_number`). | | `redactPiiReturnUnredacted` | boolean | Return the original unredacted transcript alongside the redacted one. Requires `redactPii`. | | `redactPiiSub` | enum | Substitution method for redacted PII. Allowed values: `entity_name`, `hash`. | | `redactStaticEntities` | object | Map of user-defined labels to exact terms to redact, e.g. `{ INTERNAL_TOOL: ['Bearclaw'] }`. Applied on top of standard PII redaction. Requires `redactPii`. | | `removeAudioTags` | enum | Remove inline annotations from rich transcripts. Allowed values: `all`, `speaker`. Universal-3 Pro models. | | `sentimentAnalysis` | boolean | Perform sentiment analysis on the transcript. | | `speakerLabels` | boolean | Label different speakers in the transcript (diarization). | | `speakerOptions` | object | Diarization options: `minSpeakersExpected` (number), `maxSpeakersExpected` (number). | | `speakersExpected` | number | Expected number of speakers in the audio. | | `speechThreshold` | number | Threshold for speech detection (0-1). | | `summarization` | boolean | Generate a summary of the transcript. | | `summaryModel` | enum | Summarization model. Allowed values: `informative`, `conversational`, `catchy`. | | `summaryType` | enum | Type of summary. Allowed values: `bullets`, `bullets_verbose`, `gist`, `headline`, `paragraph`. | | `temperature` | number | Sampling temperature (0-1) controlling randomness. Universal-3 Pro models. | | `webhookAuthHeaderName` | string | Name of the authentication header for webhook requests. | | `webhookAuthHeaderValue` | string | Value of the authentication header for webhook requests. | | `webhookUrl` | string | URL to send webhook notifications to. | | `wordBoost` | array of strings | Words to boost in the transcript. **Deprecated** — rejected by `universal-3-pro`, `universal-3-5-pro`, and `slam-1` (works only on `universal-2`/`best`); use `keytermsPrompt` instead. | ## Getting the full results back `transcribe()` returns a provider-agnostic result, but AssemblyAI's richer output — utterances, entities, sentiment, and more — is preserved, so you don't lose anything by going through the AI SDK: - **Top-level fields** like `result.text`, `result.segments`, and `result.durationInSeconds` are normalized across providers. - **`result.providerMetadata.assemblyai`** holds AssemblyAI-specific results such as `utterances`, `entities`, and `sentimentAnalysisResults`. - **`result.response.body`** is the complete, raw AssemblyAI transcript object. ```typescript const result = await transcribe({ model: assemblyai.transcription('universal-3-5-pro'), audio: await readFile('meeting.mp3'), providerOptions: { assemblyai: { speakerLabels: true, sentimentAnalysis: true }, }, }); // Normalized, provider-agnostic fields console.log(result.text); console.log(result.segments); // AssemblyAI-specific results const { utterances, sentimentAnalysisResults } = result.providerMetadata.assemblyai; // The complete, raw AssemblyAI transcript const rawTranscript = result.response.body; ``` Timestamps use **different units depending on where you read them.** The top-level `result.segments` are in **seconds** (the AI SDK's normalized format), while every timing inside `result.providerMetadata.assemblyai` and `result.response.body` is in **milliseconds** (AssemblyAI's native format). Convert accordingly when you combine the two. ## Additional resources - [AssemblyAI provider reference on ai-sdk.dev](https://ai-sdk.dev/providers/ai-sdk-providers/assemblyai) — the canonical, complete list of provider options - [`vercel/ai` on GitHub](https://github.com/vercel/ai) — the AI SDK source and where the provider is maintained - [AssemblyAI API reference](/api-reference/overview) — the underlying transcription API and every parameter it accepts --- # Integrate Power Automate with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/power-automate Source: docs/integrations/power-automate.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Use our Power Automate & Azure Logic Apps connector to use AssemblyAI's speech AI in your flows. [Microsoft Power Automate](https://www.microsoft.com/en-us/power-platform/products/power-automate) is a low-code workflow automation platform with a rich collection of connectors to Microsoft's first-party services and third-party services. [Azure Logic Apps](https://learn.microsoft.com/en-us/azure/logic-apps/logic-apps-overview) is the equivalent service built for developers and IT pros. The AssemblyAI connector makes our API available to both Microsoft Power Automate and Azure Logic Apps. With the connector, you can use AssemblyAI to transcribe audio data with speech recognition models, analyze the data with speech understanding models, and build generative features on top of it with LLMs. You can supply audio to the AssemblyAI connector and connect the output of our models to other services in your flows. ## Quickstart Create or edit a flow in Power Automate. Add a new action, search for AssemblyAI, and select the action that you want to use. ![Search for AssemblyAI actions in Power Automate](/assets/img/integrations/power-automate/search-action.png) You will be prompted to create a connection to AssemblyAI. Give your connection a name and enter the API key from your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home), and click **Create new**. ![Create a connection to AssemblyAI in Power Automate](/assets/img/integrations/power-automate/create-connection.png) Finally, configure your AssemblyAI action. Continue reading to learn more about all the available actions. ![Configure an AssemblyAI action in Power Automate](/assets/img/integrations/power-automate/configure-action.png) ## Upload a File To transcribe an audio file using AssemblyAI, the file needs to be accessible to AssemblyAI. If your audio file is already accessible via a URL, you can use your existing URL. Otherwise, you can use the `Upload a Media File` action to upload a file to AssemblyAI. You will get back a URL for your file which can only be used to transcribe using your API key. Once you transcribe the file, the file will be removed from AssemblyAI's servers. ## Transcribe Audio To transcribe your audio, configure the `Audio URL` parameter using your audio file URL. Then, configure the additional parameters to enable more [Speech Recognition](/pre-recorded-audio) features and [Speech Understanding](/speech-understanding) models. The result of the Transcribe Audio action is a transcript that will start being processed immediately. To get the completed transcript, you have two options: 1. [Handle the Transcript Ready Webhook](#handle-the-transcript-ready-webhook) 2. [Poll the Transcript Status](#poll-the-transcript-status) ### Handle the Transcript Ready Webhook If you don't want to handle the webhook using Logic Apps or Power Automate, configure the `Webhook URL` parameter in your `Transcribe Audio` action, and implement your webhook following [AssemblyAI's webhook documentation](/pre-recorded-audio/webhooks#handle-webhook-deliveries). To handle the webhook using Logic Apps or Power Automate, follow these steps: Create a separate Logic App or Power Automate Flow. Configure `When an HTTP request is received` as the trigger: - Set `Who Can Trigger The Flow?` to `Anyone` - Set `Request Body JSON Schema` to: ```json { "type": "object", "properties": { "transcript_id": { "type": "string" }, "status": { "type": "string" } } } ``` - Set `Method` to `POST` Add an AssemblyAI `Get Transcript` action, passing in the `transcript_id` from the trigger to the `Transcript ID` parameter. Before doing anything else, you should check whether the `Status` is `completed` or `error`. Add a `Condition` action that checks if the `Status` from the `Get Transcript` output is `error`: - In the `True` branch, add a `Terminate` action - Set the `Status` to `Failed` - Set the `Code` to `Transcript Error` - Pass the `Error` from the `Get Transcript` output to the `Message` parameter. - You can leave the `False` branch empty. Now you can add any action after the `Condition` knowing the transcript status is `completed`, and you can retrieve any of the output properties of the `Get Transcript` action. Save your Logic App or Flow. The `HTTP URL` will be generated for the `When an HTTP request is received` trigger. Copy the `HTTP URL` and head back to your original Logic App or Flow. In your original Logic App or Flow, update the `Transcribe Audio` action. Paste the `HTTP URL` you copied previously into the `Webhook URL` parameter, and save. When the transcript status becomes `completed` or `error`, AssemblyAI will send an HTTP POST request to the webhook URL, which will be handled by your other Logic App or Flow. As an alternative to using the webhook, you can poll the transcript status as explained in the next section. ### Poll the Transcript Status You can poll the transcript status using the following steps: Add an `Initialize variable` action - Set `Name` to `transcript_status` - Set `Type` to `String` - Store the `Status` from the `Transcribe Audio` output into the `Value` parameter Add a `Do until` action - Configure the `Loop Until` parameter with the following Fx code: ```plaintext or(equals(variables('transcript_status'), 'completed'), equals(variables('transcript_status'), 'error')) ``` This code checks whether the `transcript_status` variable is `completed` or `error`. - Configure the `Count` parameter to `86400` - Configure the `Timeout` parameter to `PT24H` Inside the `Do until` action, add the following actions: - Add a `Delay` action that waits for one second - Add a `Get Transcript` action and pass the `ID` from the `Transcribe Audio` output to the `Transcript ID` parameter - Add a `Set variable` action - Set `Name` to `transcript_status` - Pass the `Status` of the `Get Transcript` output to the `Value` parameter The `Do until` loop will continue until the transcript is completed, or an error occurred. Add another `Get Transcript` action, like before, but add it after the `Do until` loop so its output becomes available outside the scope of the `Do until` action. Before doing anything else, you should check whether the transcript `Status` is `completed` or `error`. Add a `Condition` action that checks if the `transcript_status` is `error`: - In the `True` branch, add a `Terminate` action - Set `Status` to `Failed` - Set `Code` to `Transcript Error` - Pass the `Error` from the `Get Transcript` output to the `Message` parameter. - You can leave the `False` branch empty. Now you can add any action after the `Condition` knowing the transcript status is `completed`, and you can retrieve any of the output properties of the `Get Transcript` action. ## Connector actions The AssemblyAI app for Power Automate provides the following actions: ### Files #### Upload a Media File Upload a media file to AssemblyAI's servers. You can pass the `Upload URL` output field to the `Audio URL` input field of [Transcribe an Audio File](#transcribe-audio) action. ### Transcripts #### Transcribe Audio Create a transcript from a media file that is accessible via a URL. Configure the `Audio URL` field with the URL of the audio file you want to transcribe. The `Audio URL` must be accessible by AssemblyAI's servers. If you don't have a publicly accessible URL, you can use the [Upload a File](#upload-a-file) action to upload the audio file to AssemblyAI. **Wait until transcript is ready** The output of this action is a transcript that is not yet `completed`. Learn [how to wait until the transcript is ready here](#transcribe-audio). Configure your desired [Speech Understanding models](/speech-understanding) when you create the transcript. The results of the models will be included in the transcript output when the transcript is completed. #### Get Transcript Get the transcript resource. The transcript is ready when the `status` is `completed`. #### Get Paragraphs in Transcript Get the transcript split by paragraphs. The API semantically segments your transcript into paragraphs to create more reader-friendly transcripts. You can only invoke this action after the transcript is completed. #### Get Sentences in Transcript Get the transcript split by sentences. The API semantically segments the transcript into sentences to create more reader-friendly transcripts. You can only invoke this action after the transcript is completed. #### Get Subtitles for Transcript Get the transcript resource. The transcript is ready when the `status` is `completed`. You can only invoke this action after the transcript is completed. #### Get Redacted Audio First, you need to configure PII audio redaction using these fields when you create the transcript: - `Redact PII`: `Yes` - `Redact PII Audio`: `Yes` - `Redact PII Policies`: Configure at least one PII policy Then, you can use this action to retrieve the redacted audio of the transcript. You can only invoke this action after the transcript is completed. #### Search Words in Transcript Search through the transcript for keywords. You can search for individual words, numbers, or phrases containing up to five words or numbers. You can only invoke this action after the transcript is completed. #### List Transcripts Get the transcript resource. The transcript is ready when the `status` is `completed`. #### Delete Transcript Delete the transcript. Deleting does not delete the resource itself, but removes the data from the resource and marks it as deleted. You can only invoke this action after the transcript status is `completed` or `error`. ## Additional resources You can learn more about using Power Automate with AssemblyAI in these resources: - [Power Automate & Logic Apps docs by Microsoft](https://learn.microsoft.com/en-us/connectors/assemblyai/) --- # Semantic Kernel Integration for AssemblyAI URL: https://www.assemblyai.com/docs/integrations/semantic-kernel Source: docs/integrations/semantic-kernel.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Transcribe audio in Semantic Kernel for C# .NET using the built-in integration with AssemblyAI. Semantic Kernel is an SDK for multiple programming languages to develop applications with [Large Language Models (LLMs)](https://www.assemblyai.com/blog/introduction-large-language-models-generative-ai/#what-are-language-models). However, LLMs only operate on textual data and don't understand what is said in audio files. With the [AssemblyAI integration for Semantic Kernel](https://github.com/AssemblyAI/assemblyai-semantic-kernel), you can use AssemblyAI's transcription models using the `TranscribePlugin` to transcribe your audio and video files. ## Quickstart Add the [AssemblyAI.SemanticKernel NuGet package](https://www.nuget.org/packages/AssemblyAI.SemanticKernel) to your project. ```bash dotnet add package AssemblyAI.SemanticKernel ``` ```powershell Install-Package AssemblyAI.SemanticKernel ``` Next, register the `TranscriptPlugin` into your kernel: ```csharp using AssemblyAI.SemanticKernel; using Microsoft.SemanticKernel; // Build your kernel var kernel = Kernel.CreateBuilder(); // Get AssemblyAI API key from env variables, or much better, from .NET configuration string apiKey = Environment.GetEnvironmentVariable("ASSEMBLYAI_API_KEY") ?? throw new Exception("ASSEMBLYAI_API_KEY env variable not configured."); kernel.ImportPluginFromObject( new TranscriptPlugin(apiKey: apiKey) ); ``` ## Usage Get the `Transcribe` function from the transcript plugin and invoke it with the context variables. ```csharp var result = await kernel.InvokeAsync( nameof(TranscriptPlugin), TranscriptPlugin.TranscribeFunctionName, new KernelArguments { ["INPUT"] = "https://assembly.ai/espn.m4a" } ); Console.WriteLine(result.GetValue()); ``` You can get the transcript using `result.GetValue()`. You can also upload local audio and video file. To do this: - Set the `TranscriptPlugin.AllowFileSystemAccess` property to `true`. - Configure the `INPUT` variable with a local file path. ```csharp kernel.ImportPluginFromObject( new TranscriptPlugin(apiKey: apiKey) { AllowFileSystemAccess = true } ); var result = await kernel.InvokeAsync( nameof(TranscriptPlugin), TranscriptPlugin.TranscribeFunctionName, new KernelArguments { ["INPUT"] = "https://assembly.ai/espn.m4a" } ); Console.WriteLine(result.GetValue()); ``` You can also invoke the function from within a semantic function like this. ```csharp string prompt = """ Here is a transcript: {{TranscriptPlugin.Transcribe "https://assembly.ai/espn.m4a"}} --- Summarize the transcript. """; var result = await kernel.InvokePromptAsync(prompt); Console.WriteLine(result.GetValue()); ``` ## Additional resources You can learn more about using Semantic Kernel with AssemblyAI in these resources: - [Ask .NET Rocks! questions with Semantic Kernel, GPT, and Chroma DB](https://www.assemblyai.com/blog/ask-dotnetrocks-questions-semantic-kernel/) - [AssemblyAI integration for Semantic Kernel GitHub repository](https://github.com/AssemblyAI/assemblyai-semantic-kernel) --- # Integrate Activepieces with AssemblyAI URL: https://www.assemblyai.com/docs/integrations/activepieces Source: docs/integrations/activepieces.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Add Speech AI to your Activepieces flows with the AssemblyAI piece. [Activepieces](https://www.activepieces.com/) is an open-source, no-code automation platform that enables users to streamline workflows by connecting various applications and automating tasks. With the AssemblyAI piece for Activepieces, you can use AssemblyAI to transcribe audio data with speech recognition models, analyze the data with speech understanding models, and build generative features on top of it with LLMs. You can supply audio to the AssemblyAI piece and connect the output of any of AssemblyAI's models to other services in your Activepieces flow. ## Quickstart Create or edit a flow in Activepiece. Add a trigger of your choosing, and then click the plus-icon to add a new action. Search for AssemblyAI, click on the AssemblyAI piece, and select the action that you want to use. ![Add an AssemblyAI piece action](/assets/img/integrations/activepieces/add-action.png) Create a new connection or select an existing one. ![Select a connection to AssemblyAI in Activepieces](/assets/img/integrations/activepieces/select-api-key.png) In the **API Key** field, enter the API key from your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home), and click **Save**. ![Configure your connection to AssemblyAI in Activepieces](/assets/img/integrations/activepieces/configure-api-key.png) Finally, configure your AssemblyAI action. Continue reading to learn more about all the available action. ## AssemblyAI actions The AssemblyAI piece for Activepieces provides the following actions: ### Files #### Upload File Upload an audio file to AssemblyAI so you can transcribe it. You can pass the `Upload URL` output field to the `Audio URL` input field of Transcribe an Audio File module. ### Transcripts #### Transcribe Transcribe an audio file and wait until the transcript has completed or failed. Configure the `Audio URL` field with the URL of the audio file you want to transcribe. The `Audio URL` must be accessible by AssemblyAI's servers. If you don't have a publicly accessible URL, you can use the Upload a File module to upload the audio file to AssemblyAI. If you don't want to wait until the transcript is ready, uncheck the `Wait until transcript is ready` parameter. Configure your desired [Speech Understanding models](/speech-understanding) when you create the transcript. The results of the models will be included in the transcript output. #### Get Transcript Retrieve a transcript by ID. #### Get Transcript Paragraphs Retrieve the paragraphs of a transcript. You can only invoke this module after the transcript is completed. #### Get Transcript Sentences Retrieve the sentences of a transcript. You can only invoke this module after the transcript is completed. #### Get Transcript Subtitles Create SRT or VTT subtitles for a transcript. You can only invoke this module after the transcript is completed. #### Get Transcript Redacted Audio First, you need to configure PII audio redaction using these fields when you create the transcript: - `Redact PII`: `Checked` - `Redact PII Audio`: `Checked` - `Redact PII Policies`: Configure at least one PII policy Then, you can use this module to retrieve the redacted audio of the transcript. You can only invoke this module after the transcript is completed. #### Search words in transcript Search for words in a transcript. You can only invoke this module after the transcript is completed. #### List transcripts Paginate over all transcripts. #### Delete transcript Delete a transcript by ID. Deleting a transcript doesn't delete the transcript resource itself, but removes the data from the resource and marks it as deleted. You can only invoke this module after the transcript status is "completed" or "error". ### Other actions #### Custom API Call Make your own REST API HTTP requests to the AssemblyAI API using your existing connection. ## Additional resources You can learn more about using Activepieces with AssemblyAI in these resources: - [AssemblyAI Integrations on Activepieces](https://www.activepieces.com/pieces/assemblyai) - [npmjs page for @activepieces/piece-assemblyai](https://www.npmjs.com/package/@activepieces/piece-assemblyai) --- # Haystack Integration for AssemblyAI URL: https://www.assemblyai.com/docs/integrations/haystack Source: docs/integrations/haystack.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Transcribe, summarize and diarize audio in a Haystack pipeline with Python using the integration with AssemblyAI. [Haystack (2.x)](https://github.com/deepset-ai/haystack) is an open-source Python framework for building custom LLM applications. The Haystack Integration for AssemblyAI seamlessly integrates with Haystack to use audio files in LLM pipelines. On top of audio transcription the AssemblyAITranscriber offers summarization and speaker diarization. This makes it possible to not only convert audio to text but also obtain concise summaries and identify speakers in a conversation. ## Quickstart Install the [assemblyai-haystack package](https://pypi.org/project/assemblyai-haystack/) using pip. This package installs and uses the AssemblyAI Python SDK and Haystack 2.0. You can find more information about the SDK at the [AssemblyAI Python SDK GitHub repository](https://github.com/AssemblyAI/assemblyai-python-sdk). ```bash pip install assemblyai-haystack ``` ## Usage Add an `AssemblyAITranscriber` component and initialize it by passing your AssemblyAI API key. Once the pipeline is ready to run, make sure to pass at least the `file_path` argument to the `run` function. The `file_path` can be a `URL` or a local file path. In the `run` function, you can also specify whether you want summarization and speaker diarization results. ```python expandable import os from assemblyai_haystack.transcriber import AssemblyAITranscriber from haystack import Pipeline from haystack.components.writers import DocumentWriter from haystack.document_stores.in_memory import InMemoryDocumentStore ASSEMBLYAI_API_KEY = os.environ.get("ASSEMBLYAI_API_KEY") document_store = InMemoryDocumentStore() file_url = "https://assembly.ai/wildfires.mp3" indexing = Pipeline() indexing.add_component("transcriber", AssemblyAITranscriber(api_key=ASSEMBLYAI_API_KEY)) indexing.add_component("writer", DocumentWriter(document_store)) indexing.connect("transcriber.transcription", "writer.documents") indexing.run( { "transcriber": { "file_path": file_url, "summarization": None, "speaker_labels": None, } } ) print("Indexed Document Count:", document_store.count_documents()) ``` Calling `indexing.run()` blocks until the transcription is finished. The results of the transcription, summarization and speaker diarization are returned in separate document lists: - `transcription` - `summarization` - `speaker_labels` When `AssemblyAITranscriber` is used in a Haystack pipeline, transcription happens by default. In the metadata of the transcription, you will also get the `ID` of the transcription and the `URL` of your audio file. A bullet point summary of what is being discussed will be returned if `summarization` is set to `TRUE`. The transcription divided into utterances of speakers will be returned if `speaker_labels` is set to `TRUE`. The output of the `AssemblyAITranscriber` is a Haystack document. When all features are turned on, the created document looks like this: ```python expandable { "transcription": [Document( id=bdf3eb20f6440cf4b15fa4fa3176eeb72bf0139a3ad4c76741724132907a5daa, content: "Smoke from hundreds of wildfires in Canada is triggering air quality alerts throughout the US. Skyli...", meta: { 'transcript_id': '2335cc07-1fbf-48ba-9855-7db3eeeb80f4', 'audio_url': "https://assembly.ai/wildfires.mp3" } ) ], "summarization": [Document( id=f88864d9229b30013d5248156e74d5bfd4435e73aadb0c0ce79040be10a4f308, content: "- Smoke from hundreds of wildfires in Canada is triggering air quality alerts...")], "speaker_labels": [Document( id=a7e222bc6a965ab1032401a6fa22da2e774294ce049b9d228acbb8b100ea2ecf, content: "Smoke from hundreds of wildfires in Canada is triggering air quality...", meta: { 'speaker': 'A' } ), Document( id=711a1888af58601e6392490a5e4ca4c10958f93a52d8f0734869c54573ea76f5, content: "Well, there's a couple of things. The season has been pretty dry already...", meta: { 'speaker': 'B' } ), Document( id=8fc78631d420e2e6127b8bdff2830f693febb91ed1566b9a84527cf023023d9e, content: "So what is it in this haze that makes it harmful?", meta: { 'speaker': 'A' } ), ... ]} ``` ## Additional resources You can learn more about using Haystack with AssemblyAI in these resources: - [Announcing the AssemblyAI Integration for Haystack](https://www.assemblyai.com/blog/announcing-the-assemblyai-integration-for-haystack/) - [AssemblyAI integration for Haystack GitHub repository](https://github.com/AssemblyAI/assemblyai-haystack) --- # Cloudflare URL: https://www.assemblyai.com/docs/nav-links/cloudflare Source: docs/nav-links/cloudflare.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Open Cloudflare. --- # Relay.app URL: https://www.assemblyai.com/docs/nav-links/relay-app Source: docs/nav-links/relay-app.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Open Relay.app. --- # Bubble by Knowcode URL: https://www.assemblyai.com/docs/nav-links/bubble-by-knowcode Source: docs/nav-links/bubble-by-knowcode.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Open Bubble by Knowcode. --- # Pipedream URL: https://www.assemblyai.com/docs/nav-links/pipedream Source: docs/nav-links/pipedream.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Open Pipedream. --- # Drupal URL: https://www.assemblyai.com/docs/nav-links/drupal Source: docs/nav-links/drupal.mdx Navigation: Overview > Use cases & integrations > Integrations > Community-maintained tools Description: Open Drupal. --- # Trust center URL: https://www.assemblyai.com/docs/nav-links/trust-center Source: docs/nav-links/trust-center.mdx Navigation: Overview > Trust & security Description: Open Trust center. --- # Security overview URL: https://www.assemblyai.com/docs/nav-links/security-overview Source: docs/nav-links/security-overview.mdx Navigation: Overview > Trust & security Description: Open Security overview. --- # Data Controls URL: https://www.assemblyai.com/docs/data-controls Source: docs/data-controls.mdx Navigation: Overview > Trust & security Description: Manage how AssemblyAI retains and uses your data — opt out of the model improvement program, set a time-to-live for audio and transcripts, and sign a BAA, all self-serve from the dashboard. The [**Data Controls** page](https://www.assemblyai.com/dashboard/settings/data-controls) in the AssemblyAI dashboard lets you manage how AssemblyAI retains and uses your data. From a single page you can opt out of the model improvement program, set a time-to-live (TTL) for your audio and transcripts, and review and sign a Business Associate Agreement (BAA). The page always reflects the **current state of your account** so you can see, at a glance, your model-improvement opt-out status, your active data-retention (TTL) setting, and whether a BAA is in place. All three of these controls — opt out, TTL, and BAA — are available **self-serve at no additional cost**. You can manage them directly from the dashboard without contacting sales or support. ## Who can access Data Controls You must be on a **paid plan** to access Data Controls. Free users cannot opt out of the model improvement program, set a TTL, or sign a BAA. To access these settings, [upgrade to a paid plan](https://www.assemblyai.com/dashboard/pricing). Only members with the **Owner** or **Admin** role can view and change Data Controls settings. Members with the **Reader** role do not have access to this page. For more on roles and permissions, see [Account Management](/account-management#roles). ## Opt out of the model improvement program Toggle **Opt Out of Data Sharing for Model Improvement Program** to control whether AssemblyAI may use your Customer Data to train its models. When you opt out, at this time: - AssemblyAI will not use your Customer Data to train its artificial intelligence and machine learning models. - AssemblyAI will not use your Customer Data to perform benchmarking. - AssemblyAI will not use your Deidentified Data to train its artificial intelligence and machine learning models. Opt-out changes are forward-looking only and apply to subsequent new requests. For full details on how model training works, see [Data retention and model training](/data-retention-and-model-training#model-training). ## Set a time-to-live for audio and transcripts The **time-to-live (TTL)** mechanism controls how long AssemblyAI retains your audio and transcripts in the asynchronous production environment. When you set a TTL, AssemblyAI begins the deletion process for your audio and transcripts at the set TTL time. **This TTL also applies to the inputs and outputs of LLM Gateway.** Certain metadata is stored for logging and billing purposes. You can choose a preset (1 day, 3 days, 7 days, or 30 days) or set a **Custom** value. Any changes are applied to subsequent new requests. For more detail, including information on potential TTL deletion lag times, see [Data retention and model training](/data-retention-and-model-training#data-retention). ## Business Associate Agreement (BAA) AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). If you need a BAA in place to process PHI, you can review and sign AssemblyAI's standard online BAA terms directly from the Data Controls page. The page shows your current BAA status (for example, **Not signed**). To execute a BAA, select **Review & sign BAA** and complete the online signing flow. You must have an upgraded (paid) account to see the option to initiate a BAA in the dashboard. If you're on the free plan, the BAA controls are hidden — [upgrade](https://www.assemblyai.com/dashboard/pricing) first. ### How signing a BAA affects your opt-out status Once you have signed a BAA, you are **automatically opted out** of the model improvement program upon signature, and the **opt-out toggle can no longer be changed**. AssemblyAI does not use files submitted under a BAA for model training. Signing a BAA also changes your default data retention behavior. For BAA-specific retention timelines, see the [asynchronous production environment](/data-retention-and-model-training#asynchronous-production-environment) and [LLM Gateway](/data-retention-and-model-training#llm-gateway-production-environment) retention tables. ## Related resources Learn how AssemblyAI handles data retention, encryption, model training, and compliance. Understand roles and permissions, including who can access Data Controls. --- # Data retention and model training URL: https://www.assemblyai.com/docs/data-retention-and-model-training Source: docs/data-retention-and-model-training.mdx Navigation: Overview > Trust & security Description: Learn about how AssemblyAI handles data retention, encryption, model training, and compliance. ## Model training We consider model training critical to providing you with the most accurate models and services that we can. Only certain files submitted to the API, as permitted by the applicable contract, are used for model training. These files undergo a redaction process designed to redact personally identifiable information before any remaining data is used for model training. We will not use files you submit for model training if you are subject to a Business Associate Addendum, are utilizing [our European servers](/pre-recorded-audio/select-the-region), or if you have opted out from model training. You can find more information on if and how to opt out [here](/faq/how-to-opt-out-of-data-sharing-for-our-model-improvement-program). ## LLM Gateway model training AssemblyAI has opted out of data training with all LLM Gateway providers. Please note this is separate from whether AssemblyAI may train our models with your data. You can find more information on if and how to opt out of data sharing for our model improvement program [here](/faq/how-to-opt-out-of-data-sharing-for-our-model-improvement-program). ## Encryption Data at rest is encrypted with AES 128 or AES-256, and data in transit uses TLS 1.2+. AssemblyAI posts SSL scans quarterly to its [Trust Center](https://app.vanta.com/assemblyai/trust/7n80syl8zln1bn1qm3x8eg) to verify the use of TLS with modern ciphersuites to its service. ### Async For transcription of pre-recorded audio, AssemblyAI supports the following TLS versions and cipher suites: **Supported TLS versions:** - TLS 1.3 - TLS 1.2 **Supported cipher suites:** - TLS_AES_128_GCM_SHA256 - TLS_AES_256_GCM_SHA384 - TLS_CHACHA20_POLY1305_SHA256 - ECDHE-ECDSA-AES128-GCM-SHA256 - ECDHE-RSA-AES128-GCM-SHA256 - ECDHE-ECDSA-AES256-GCM-SHA384 - ECDHE-RSA-AES256-GCM-SHA384 ### Streaming For transcription of streaming audio, AssemblyAI supports the following TLS version and cipher suites: **Supported TLS versions:** - TLS 1.3 **Supported cipher suites:** - TLS_AES_128_GCM_SHA256 - TLS_AES_256_GCM_SHA384 - TLS_CHACHA20_POLY1305_SHA256 Ensure your client or application is configured to use one of the supported TLS versions and cipher suites when connecting to AssemblyAI services. ## GDPR compliance We have designed our products with GDPR principles top of mind but also understand that privacy compliance is a moving target. As privacy requirements continue to evolve (rapidly), we are constantly working to assess and improve our practices. You can read more about our privacy practices in our Privacy Policy [here](https://www.assemblyai.com/legal/privacy-policy), and Data Processing Addendum [here](https://www.assemblyai.com/legal/data-processing-addendum). ## SOC2 certification We have both SOC2 Type 1 and Type 2 certifications. You can find more information on this on our [Trust Center](https://www.assemblyai.com/trust). We also have a great blog post on the subject, which you can find [here](https://www.assemblyai.com/blog/assemblyai-obtains-soc2-type-2-compliance-for-2022-2023/). ## Data retention ### Streaming production environment If you are opted out of model training, we offer zero data retention of audio and transcripts for our Streaming product. Certain metadata about the transcript is stored and maintained for logging and billing purposes. The model training environment differs from the production environment. You can find more information on model training in our [Model Training section](#model-training). If you would like to opt out of model training, please see our [Opt-Out FAQ](/faq/how-to-opt-out-of-data-sharing-for-our-model-improvement-program). ### Asynchronous production environment | Artifact Type\* | Time-To-Live Configured | BAA Executed | No TTL or BAA | Customer-Initiated Deletion Request | | --------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | ------------------------------------------------- | | **Customer-Uploaded Audio Files** | Deletion process begins in AWS at TTL expiration, which can be as low as one (1) hour, subject to AWS TTL processing times\*\*

If TTL is longer than 24 hours, deletion process will automatically begin at 24 hours and is at most 48 hours. | Default deletion process begins at 72 hours (or can be set to as low as 1-hour), subject to processing times due to the AWS TTL\*\* | Deletion process begins at 24 hours and is at most 48 hours. | Deleted when customer initiates deletion request. | | **Final Transcription Artifact** | Deletion process begins in AWS at TTL expiration, which can be as low as one (1) hour, subject to AWS TTL processing times\*\* | Default deletion process begins at 72 hours (or can be set to as low as 1-hour), subject to processing times due to the AWS TTL\*\* | Deletion process begins at 30 days and is automatic thereafter. | Deleted when customer initiates deletion request. | | **Customer-Provided URL Audio Reference** | Linked to the lifecycle of a Final Transcription Artifact: Deletion process begins in AWS at TTL expiration, which can be as low as one (1) hour, subject to AWS TTL processing times\*\* | Linked to the lifecycle of a Final Transcription Artifact: Default deletion process begins at 72 hours (or can be set to as low as 1-hour), subject to processing times due to the AWS TTL\*\* | Linked to the lifecycle of a Final Transcription Artifact: Deletion process begins at 30 days and is automatic thereafter. | Deleted when customer initiates deletion request. | | **Intermediate Artifact** | Linked to the lifecycle of a Final Transcription Artifact: Deletion process begins in AWS at TTL expiration, which can be as low as one (1) hour, subject to AWS TTL processing times\*\*

If TTL is longer than 48 hours, deletion process will automatically begin at 48 hours and is at most 72 hours. | Linked to the lifecycle of a Final Transcription Artifact: Default deletion process begins at 72 hours (or can be set to as low as 1-hour), subject to processing times due to the AWS TTL\*\* | Deletion process begins at 48 hours and is at most 72 hours. | Deleted when customer initiates deletion request. | | **Transcription Request Text Inputs** (this includes content inputs such as key terms prompts, word boost list, etc.) | Linked to the lifecycle of a Final Transcription Artifact: Deletion process begins in AWS at TTL expiration, which can be as low as one (1) hour, subject to AWS TTL processing times\*\* | Linked to the lifecycle of a Final Transcription Artifact: Default deletion process begins at 72 hours (or can be set to as low as 1-hour), subject to processing times due to the AWS TTL\*\* | Linked to the lifecycle of a Final Transcription Artifact: Deletion process begins at 30 days and is automatic thereafter. | Deleted when customer initiates deletion request. | \*Certain metadata is stored for logging and billing purposes. \*\*The minimum TTL that AssemblyAI may set for Final Transcription Artifacts in the asynchronous production environment is 1 (one) hour. The TTL mechanism that AssemblyAI uses is through Amazon Web Services's ("AWS") DynamoDB TTL mechanism (the "AWS TTL"). The deletion process begins in AWS at TTL expiration, but is subject to AWS TTL processing times. In practice, these deletion events typically take place anywhere from a few minutes to a few hours after the deletion process begins in AWS, depending on circumstances, including server location. However, we have seen lag times anywhere from 2-3 hours to a few days. Once the artifact is deleted in AWS, AssemblyAI processes this deletion almost immediately. See [here](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/TTL.html) for more information about AWS's TTL mechanism. To learn how to delete a transcript via the API, see [Delete transcripts](/pre-recorded-audio/delete-transcripts). #### Confirming deletion Should you wish to confirm a file has been deleted, or in case you did not store the `transcript_id` when the transcription request was made, you can get a list of all transcripts. You can make a GET request to `https://api.assemblyai.com/v2/transcript` which will return a list of all transcripts created or specify a `transcript_id` to review a single transcript. The model training environment differs from the production environment. You can find more information on model training in our [Model Training section](#model-training). If you would like to opt out of model training, please see our [Opt-Out FAQ](/faq/how-to-opt-out-of-data-sharing-for-our-model-improvement-program). ### LLM Gateway production environment - If you have an executed BAA and use either Anthropic or Google inference models, we offer zero data retention for LLM Gateway inputs and outputs. Certain metadata is stored for logging and billing purposes. - If you have a designated TTL on your LLM Gateway account, we delete inputs and outputs on an hourly basis. Certain metadata is stored for logging and billing purposes. - If a customer initiates a deletion request, inputs and outputs are deleted at the time of the request. Certain metadata is stored for logging and billing purposes. For deletion of speech understanding requests, please see below. - For speech understanding requests, such as translation, speaker ID, or custom formatting, the retention is linked to the life of an asynchronous Final Transcription Artifact, noted above. #### Provider-specific retention policies **Anthropic Claude models** We use Anthropic models through Amazon Bedrock. Amazon Bedrock doesn't store or log your prompts and completions. Amazon Bedrock doesn't use your prompts and completions to train any AWS models and doesn't distribute them to third parties. See [here](https://docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html) for more information on Amazon Bedrock data protection policies. If Amazon Bedrock fails, for non-EU customers we may send your request to the Anthropic API, where we have 0-day retention configured. Please see Anthropic's commercial terms [here](https://www.anthropic.com/legal/commercial-terms). **OpenAI GPT models** OpenAI models have ZDR (zero data retention). See OpenAI's policy [here](https://developers.openai.com/api/docs/guides/your-data#zero-data-retention) for more information on how OpenAI defines ZDR. For OpenAI open-weight models (gpt-oss-120b, gpt-oss-20b), we use these models through Amazon Bedrock. Amazon Bedrock doesn't store or log your prompts and completions. Amazon Bedrock doesn't use your prompts and completions to train any AWS models and doesn't distribute them to third parties. See [here](https://docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html) for more information on Amazon Bedrock data protection policies. **Google Gemini models** Google Gemini models have ZDR (zero data retention). See Google's policy [here](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/vertex-ai-zero-data-retention) for more information on how Google defines ZDR. #### Opt-in providers The following providers are opt-in only. To use them with Speech Understanding or the LLM Gateway, opt in from the [Data Controls](/data-controls) page in your dashboard. The providers listed above are enabled by default when you use Speech Understanding or the LLM Gateway with those models by name. **Together AI** Together AI has ZDR (zero data retention) enabled. See Together AI's [privacy policy](https://www.together.ai/privacy) for more information. **Fireworks AI** Fireworks AI has [ZDR by default](https://docs.fireworks.ai/guides/security_compliance/data_handling#zero-data-retention). See Fireworks AI's [data handling documentation](https://docs.fireworks.ai/guides/security_compliance/data_handling#zero-data-retention) for more information. **DigitalOcean** See DigitalOcean's [inference data privacy documentation](https://docs.digitalocean.com/products/inference/details/data-privacy/) for more information. AssemblyAI has opted out of model training with all LLM Gateway providers. Please note this is separate from whether AssemblyAI may train our models with your data, and the model training environment differs from the production environment. You can find more information on model training in our [Model Training section](#model-training). If you would like to opt out of model training, please see our [Opt-Out FAQ](/faq/how-to-opt-out-of-data-sharing-for-our-model-improvement-program). --- # API Reference URL: https://www.assemblyai.com/docs/api-reference/overview Source: docs/api-reference/overview.mdx Navigation: API Reference > Overview Description: Choose an API to view its full reference. Pick a product to jump to its full API reference. Each product tab has its own reference section with all endpoints, request and response schemas, and example calls. Get clean, customizable transcripts in 99 languages with industry-leading accuracy and natural language prompting. Real-time transcription your notetaker, agents, and captions can depend on. Send an audio file in a single HTTP request, get a transcript back in milliseconds. No polling, no session management. One OpenAI-compatible API for every frontier model — with automatic fallbacks, zero markup, and zero data retention. Stream audio in, get audio back. We handle the rest so you can focus on your product. --- # Quickstart URL: https://www.assemblyai.com/docs/pre-recorded-audio/getting-started/transcribe-an-audio-file Source: docs/pre-recorded-audio/getting-started/transcribe-an-audio-file.mdx Navigation: Pre-recorded STT > Getting started Description: Learn how to transcribe and analyze an audio file. ## Overview By the end of this guide, you'll have a working script that transcribes an audio file in a single SDK call. Build it with an AI coding agent, or write it yourself — both are below. Prefer to try it first? Transcribe audio without writing any code in the [AssemblyAI Playground](https://www.assemblyai.com/dashboard/playground/transcript). ## Before you begin You'll need: - **An API key** — grab one from [your dashboard](https://www.assemblyai.com/dashboard/home). Every example below reads it from an environment variable, so set it once: ```bash export ASSEMBLYAI_API_KEY= ``` - **Python 3.8+ or Node.js 18+**, depending on which SDK you use. **Building with an AI coding agent?** Wire it up to AssemblyAI's live docs (MCP server) and the AssemblyAI skill so it writes correct, up-to-date code instead of relying on stale training data: ```bash claude mcp add --transport http --scope user assemblyai-docs https://assemblyai.com/docs/mcp npx skills add AssemblyAI/assemblyai-skill --global ``` Then describe what you want to build. To get the same result as the steps below, paste: ```text Use the AssemblyAI Python SDK to transcribe https://assembly.ai/wildfires.mp3 and print the transcript text. ``` ## Transcribe your first file Prefer to write it yourself? Follow these steps to transcribe our hosted sample file. The SDK uploads, submits, and polls for you in a single call. ### Step 1: Install the SDK ```bash pip install assemblyai ``` ```bash npm install assemblyai ``` ### Step 2: Run your first transcription Save this as `transcribe.py` (Python) or `transcribe.js` (JavaScript): ```python import os from assemblyai.prerecorded.v2 import Transcriber transcriber = Transcriber(api_key=os.environ["ASSEMBLYAI_API_KEY"]) transcript = transcriber.transcribe("https://assembly.ai/wildfires.mp3") print(transcript.text) ``` ```javascript import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY }); const transcript = await client.transcripts.transcribe({ audio: "https://assembly.ai/wildfires.mp3", }); console.log(transcript.text); ``` Then run it — `python transcribe.py` or `node transcribe.js`. You'll see the transcript printed: ```text Smoke from hundreds of wildfires in Canada is triggering air quality alerts throughout the US... ``` That's the whole first call. From here you can add options — speaker labels, language detection, or a local file — see the [complete example](#complete-example) to combine them, or use the [HTTP API directly](#using-the-http-api-directly) if you're not using an SDK. ## Customize your request The call above works with no extra configuration. Add capabilities by setting options on the same request — combine as many as you need (the [complete example](#complete-example) sets several at once). ### Transcribe a local file Pass a file path instead of a URL; the SDK uploads it for you. ```python transcript = transcriber.transcribe("./example.mp3") ``` ```javascript const transcript = await client.transcripts.transcribe({ audio: "./example.mp3", }); ``` ### Identify speakers Enable [Speaker Diarization](/pre-recorded-audio/label-speakers) to split the transcript by speaker. Each labeled segment (an *utterance*) has a speaker ID and its text. ```python from assemblyai.prerecorded.v2 import TranscriptionConfig config = TranscriptionConfig(speaker_labels=True) transcript = transcriber.transcribe("https://assembly.ai/wildfires.mp3", config=config) for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` ```javascript const transcript = await client.transcripts.transcribe({ audio: "https://assembly.ai/wildfires.mp3", speaker_labels: true, }); for (const utterance of transcript.utterances) { console.log(`Speaker ${utterance.speaker}: ${utterance.text}`); } ``` ### Detect the language automatically Use [Automatic Language Detection](/pre-recorded-audio/language-detection) to detect the dominant spoken language. The `language_detection=True` option is used in the complete example below. ## Complete example Here's the complete, runnable script — the call above plus options and error handling: ```python expandable import os from assemblyai import TranscriptStatus from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # Use a publicly-accessible URL audio_file = "https://assembly.ai/wildfires.mp3" # Or use a local file: # audio_file = "./example.mp3" config = TranscriptionConfig( language_detection=True, speaker_labels=True, ) transcriber = Transcriber(api_key=os.environ["ASSEMBLYAI_API_KEY"]) transcript = transcriber.transcribe(audio_file, config=config) if transcript.status == TranscriptStatus.error: raise RuntimeError(f"Transcription failed: {transcript.error}") # Log transcript.id for every request (not just errors), with a timestamp and API region. # It's required to fetch results, retry, or delete the transcript later, and it's the first # thing support@assemblyai.com asks for. Delete: /pre-recorded-audio/delete-transcripts # Troubleshooting: /pre-recorded-audio/guides/common_errors_and_solutions print(f"\nFull Transcript:\n\n{transcript.text}") # Optionally print speaker diarization results # for utterance in transcript.utterances: # print(f"Speaker {utterance.speaker}: {utterance.text}") ``` ```javascript expandable import { AssemblyAI } from "assemblyai"; const baseUrl = "https://api.assemblyai.com"; const client = new AssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY, baseUrl: baseUrl, }); // Use a publicly-accessible URL const audioFile = "https://assembly.ai/wildfires.mp3"; // Or use a local file: // const audioFile = "./example.mp3"; const params = { audio: audioFile, language_detection: true, speaker_labels: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } // Log transcript.id for every request (not just errors), with a timestamp and API region. // It's required to fetch results, retry, or delete the transcript later, and it's the first // thing support@assemblyai.com asks for. Delete: /pre-recorded-audio/delete-transcripts // Troubleshooting: /pre-recorded-audio/guides/common_errors_and_solutions console.log(`\nFull Transcript:\n\n${transcript.text}`); // Optionally print speaker diarization results // for (const utterance of transcript.utterances) { // console.log(`Speaker ${utterance.speaker}: ${utterance.text}`); // } }; run(); ``` ## What you get back A completed transcript includes the full `text` plus metadata, and per-speaker `utterances` when you enable `speaker_labels`. The SDK exposes these as attributes (`transcript.text`, `transcript.utterances[0].speaker`); the raw API returns the same fields as JSON: ```json { "id": "106993b6-ac12-45d0-b74a-1bbd923e755d", "status": "completed", "text": "Smoke from hundreds of wildfires in Canada is triggering air quality alerts...", "language_code": "en", "audio_duration": 282, "confidence": 0.95, "utterances": [ { "speaker": "A", "text": "Smoke from hundreds of wildfires in Canada is triggering air quality alerts...", "confidence": 0.97, "start": 100, "end": 26560, "words": [ { "text": "Smoke", "start": 100, "end": 640, "confidence": 0.9, "speaker": "A" } ] } ] } ``` `start` and `end` are in milliseconds. Persist `id` to fetch, retry, or delete the transcript later. See the [transcript API reference](/api-reference/transcripts/get) for the complete field list. ## Using the HTTP API directly Not using an SDK? The same flow works over plain HTTP — authenticate with your key in the `authorization` header (no `Bearer` prefix), submit to `POST /v2/transcript`, then poll (repeatedly call `GET /v2/transcript/{id}`) until the status is `completed`. The SDKs above do all of this for you, including uploading local files and polling. All three examples read your key from the same `ASSEMBLYAI_API_KEY` environment variable you set in [Before you begin](#before-you-begin). The cURL example also needs [`jq`](https://jqlang.github.io/jq/) (`brew install jq`); the Python example needs the `requests` library (`pip install requests`); the JavaScript example needs Node.js 18+ (built-in `fetch`). Submit the file, poll until the status is `completed`, then print the text. (The variable is named `state` because zsh reserves `status`.) ```bash expandable id=$(curl -s -X POST https://api.assemblyai.com/v2/transcript \ -H "authorization: $ASSEMBLYAI_API_KEY" \ -H "content-type: application/json" \ -d '{ "audio_url": "https://assembly.ai/wildfires.mp3", "language_detection": true, "speaker_labels": true }' | jq -r .id) while true; do state=$(curl -s https://api.assemblyai.com/v2/transcript/$id \ -H "authorization: $ASSEMBLYAI_API_KEY" | jq -r .status) [ "$state" = "completed" ] && break [ "$state" = "error" ] && { echo "Transcription failed"; break; } sleep 3 done curl -s https://api.assemblyai.com/v2/transcript/$id \ -H "authorization: $ASSEMBLYAI_API_KEY" | jq -r .text ``` To transcribe a local file, upload it first and use the returned `upload_url` as the `audio_url`: ```bash curl -s -X POST https://api.assemblyai.com/v2/upload \ -H "authorization: $ASSEMBLYAI_API_KEY" \ --data-binary @./example.mp3 | jq -r .upload_url ``` The file must be streamed as raw bytes with `curl --data-binary @` (note the `@`). Using `-d`/`--data`, or passing a JSON body or a file-path string, will return a successful `upload_url` but then fail downstream at transcription with a `Transcoding failed. File type application/json` or `text/plain` error. See [Troubleshoot Common Errors](/pre-recorded-audio/guides/common_errors_and_solutions) for details. ```python expandable import os import requests import time base_url = "https://api.assemblyai.com" headers = {"authorization": os.environ["ASSEMBLYAI_API_KEY"]} # Use a publicly-accessible URL audio_file = "https://assembly.ai/wildfires.mp3" # Or upload a local file: # with open("./example.mp3", "rb") as f: # response = requests.post(base_url + "/v2/upload", headers=headers, data=f) # if response.status_code != 200: # print(f"Error: {response.status_code}, Response: {response.text}") # response.raise_for_status() # upload_json = response.json() # audio_file = upload_json["upload_url"] data = { "audio_url": audio_file, "language_detection": True, "speaker_labels": True } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) if response.status_code != 200: print(f"Error: {response.status_code}, Response: {response.text}") response.raise_for_status() transcript_json = response.json() transcript_id = transcript_json["id"] polling_endpoint = f"{base_url}/v2/transcript/{transcript_id}" while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": print(f"\nFull Transcript:\n\n{transcript['text']}") # Optionally print speaker diarization results # for utterance in transcript['utterances']: # print(f"Speaker {utterance['speaker']}: {utterance['text']}") break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) ``` ```javascript expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: process.env.ASSEMBLYAI_API_KEY, }; async function transcribe() { // Use a publicly-accessible URL const audioFile = "https://assembly.ai/wildfires.mp3"; // Or upload a local file: // import fs from "fs-extra"; // const audioData = await fs.readFile("./example.mp3"); // const uploadRes = await fetch(`${baseUrl}/v2/upload`, { // method: "POST", // headers, // body: audioData, // }); // if (!uploadRes.ok) throw new Error(`Error: ${uploadRes.status}`); // const uploadResponse = await uploadRes.json(); // const audioFile = uploadResponse.upload_url; const data = { audio_url: audioFile, language_detection: true, speaker_labels: true, }; let res = await fetch(`${baseUrl}/v2/transcript`, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptResponse = await res.json(); const transcriptId = transcriptResponse.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcript = await res.json(); if (transcript.status === "completed") { console.log(`\nFull Transcript:\n\n${transcript.text}`); // Optionally print speaker diarization results // for (const utterance of transcript.utterances) { // console.log(`Speaker ${utterance.speaker}: ${utterance.text}`); // } break; } else if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } } transcribe(); ``` ## Limits - **File size:** up to 5 GB per request (`/v2/transcript`); local files uploaded via `/v2/upload` up to 2.2 GB. - **Duration:** 160 ms to 10 hours per file. - **Formats:** most common audio and video formats — submit your file as-is, no transcoding needed. - **Rate limit:** default 5 parallel jobs on free accounts, 200 on paid. Check yours on the [rate limits page](https://www.assemblyai.com/dashboard/home). ## Next steps Now that you have transcribed your first audio file: - Explore [our Speech Understanding features](https://www.assemblyai.com/products/speech-understanding) for more ways to analyze your audio data - Learn more about searching, summarizing, or asking questions on your transcript with [our LLM Gateway feature](/llm-gateway/quickstart) - Find out how to use [webhooks](/pre-recorded-audio/webhooks) to get notified when your transcripts are ready For more information, check out the full [API reference documentation](/). ## Need some help? If you get stuck, or have any other questions, we'd love to help you out. Contact our support team at support@assemblyai.com or create a [support ticket](https://www.assemblyai.com/contact/support). --- # Model selection URL: https://www.assemblyai.com/docs/pre-recorded-audio/select-the-speech-model Source: docs/pre-recorded-audio/select-the-speech-model.mdx Navigation: Pre-recorded STT > Getting started Description: Model selection documentation. The `speech_models` parameter lets you specify which model(s) to use for transcription. If omitted, it defaults to `["universal-3-5-pro", "universal-2"]`. With this default, `universal-3-5-pro` handles its 18 supported languages; for all other languages, the request will automatically fall back to `universal-2`. ## Available models | Name | Parameter | Description | Best for | | ------------------- | ----------------------------------- | ------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- | | **Universal-3.5 Pro** Recommended | `speech_models=['universal-3-5-pro']` | Our highest accuracy, fastest model with 18-language support, native code switching, and contextual prompting. | Highest-accuracy transcription, post-call analytics, meeting notetakers, medical transcription, domain-specific accuracy via prompting | | **Universal-2** | `speech_models=['universal-2']` | Our accurate, cost-effective model with support across 99 languages. | High-volume batch transcription, 99-language coverage, price-sensitive workloads, fallback for unsupported U3 Pro languages | | Name | Parameter | Description | Best for | | ------------------- | ------------------------------------ | ------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- | | **Universal-3.5 Pro** Recommended | `speech_models: ['universal-3-5-pro']` | Our highest accuracy, fastest model with 18-language support, native code switching, and contextual prompting. | Highest-accuracy transcription, post-call analytics, meeting notetakers, medical transcription, domain-specific accuracy via prompting | | **Universal-2** | `speech_models: ['universal-2']` | Our accurate, cost-effective model with support across 99 languages. | High-volume batch transcription, 99-language coverage, price-sensitive workloads, fallback for unsupported U3 Pro languages | | Name | API Parameter | Description | Best for | | ------------------- | -------------------------------------- | ------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- | | **Universal-3.5 Pro** Recommended | `"speech_models": ["universal-3-5-pro"]` | Our highest accuracy, fastest model with 18-language support, native code switching, and contextual prompting. | Highest-accuracy transcription, post-call analytics, meeting notetakers, medical transcription, domain-specific accuracy via prompting | | **Universal-2** | `"speech_models":["universal-2"]` | Our accurate, cost-effective model with support across 99 languages. | High-volume batch transcription, 99-language coverage, price-sensitive workloads, fallback for unsupported U3 Pro languages | ## Set the model You can change the model by setting the `speech_models` in the POST request body: ```python highlight={12} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } data = { "audio_url": "https://assembly.ai/wildfires.mp3", "speech_models": ["universal-3-5-pro"], "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(transcription_result['text']) break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` You can change the model by setting `speech_models` in the transcription config: ```python highlight={7} from assemblyai import TranscriptStatus from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( speech_models=["universal-3-5-pro"], language_detection=True ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == TranscriptStatus.error: raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` You can change the model by setting the `speech_models` in the POST request body: ```javascript highlight={12} expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const data = { audio_url: "https://assembly.ai/wildfires.mp3", speech_models: ["universal-3-5-pro"], language_detection: true, }; const url = `${baseUrl}/v2/transcript`; let res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` You can change the model by setting the `speech_models` in the transcript parameters: ```javascript highlight={12} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, speech_models: ["universal-3-5-pro"], language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } console.log(transcript.text); }; run(); ``` ## Identify the model used After transcription completes, you can check which model was actually used to process your request by reading the `speech_model_used` field. This is useful when you provide multiple models in the `speech_models` array, as the system may fall back to a different model depending on language support. ```python expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } data = { "audio_url": "https://assembly.ai/wildfires.mp3", "speech_models": ["universal-3-5-pro"], "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Model used: {transcription_result['speech_model_used']}") print(transcription_result['text']) break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```python from assemblyai import TranscriptStatus from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( speech_models=["universal-3-5-pro"], language_detection=True ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == TranscriptStatus.error: raise RuntimeError(f"Transcription failed: {transcript.error}") print(f"Model used: {transcript.json_response['speech_model_used']}") print(transcript.text) ``` ```javascript expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const data = { audio_url: "https://assembly.ai/wildfires.mp3", speech_models: ["universal-3-5-pro"], language_detection: true, }; const url = `${baseUrl}/v2/transcript`; let res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(`Model used: ${transcriptionResult.speech_model_used}`); console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ```javascript expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, speech_models: ["universal-3-5-pro"], language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } console.log(`Model used: ${transcript.speech_model_used}`); console.log(transcript.text); }; run(); ``` ## Complete example Here is the full working code that demonstrates model selection with error handling: ```python expandable import requests import time base_url = "https://api.assemblyai.com" headers = {"authorization": ""} data = { "audio_url": "https://assembly.ai/wildfires.mp3", "speech_models": ["universal-3-5-pro"], "language_detection": True } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) if response.status_code != 200: print(f"Error: {response.status_code}, Response: {response.text}") response.raise_for_status() transcript_json = response.json() transcript_id = transcript_json["id"] polling_endpoint = f"{base_url}/v2/transcript/{transcript_id}" while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": print(f"Model used: {transcript['speech_model_used']}") print(f"\nTranscript:\n\n{transcript['text']}") break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) ``` ```python from assemblyai import TranscriptStatus from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( speech_models=["universal-3-5-pro"], language_detection=True, ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == TranscriptStatus.error: raise RuntimeError(f"Transcription failed: {transcript.error}") print(f"Model used: {transcript.json_response['speech_model_used']}") print(f"\nTranscript:\n\n{transcript.text}") ``` ```javascript expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; async function transcribe() { const audioFile = "https://assembly.ai/wildfires.mp3"; const data = { audio_url: audioFile, speech_models: ["universal-3-5-pro"], language_detection: true, }; let res = await fetch(`${baseUrl}/v2/transcript`, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptResponse = await res.json(); const transcriptId = transcriptResponse.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcript = await res.json(); if (transcript.status === "completed") { console.log(`Model used: ${transcript.speech_model_used}`); console.log(`\nTranscript:\n\n${transcript.text}`); break; } else if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } } transcribe(); ``` ```javascript expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, speech_models: ["universal-3-5-pro"], language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } console.log(`Model used: ${transcript.speech_model_used}`); console.log(`\nTranscript:\n\n${transcript.text}`); }; run(); ``` --- # Prompting and Keyterms URL: https://www.assemblyai.com/docs/pre-recorded-audio/universal-3-5-pro/prompting Source: docs/pre-recorded-audio/universal-3-5-pro/prompting.mdx Navigation: Pre-recorded STT > Features Description: Prompting and Keyterms documentation. Universal-3.5 Pro is highly accurate out of the box, but for challenging audio like short clips with limited context, noisy environments, or audio with niche references, you can give the model information about your audio to improve transcription accuracy. There are two ways to give the model information about your audio: - **Contextual prompting** (`prompt`) — a natural-language description of what the audio is about: the domain, the scenario, or the full details of the conversation. - **Keyterms prompting** (`keyterms_prompt`) — an explicit list of terms you want the model to recognize accurately. ## Contextual prompting Use the `prompt` parameter to provide context about your audio — describe what is being transcribed, not how to transcribe it. Formatting and behavioral instructions are ignored. The model stays grounded in the audio, so irrelevant context won't cause hallucinated words. For example, this is a 2-second clip from a League of Legends pro interview: Without prompt: ```txt And so look who I've been a dear. ``` With prompt: ```txt In solo queue, I ban Azir. ``` ```python {11} expandable import requests import time base_url = "https://api.assemblyai.com" headers = {"authorization": ""} data = { "audio_url": "https://assembly.ai/prompt-8", "language_detection": True, "prompt": "League of Legends roles" } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) if response.status_code != 200: print(f"Error: {response.status_code}, Response: {response.text}") response.raise_for_status() transcript_response = response.json() transcript_id = transcript_response["id"] polling_endpoint = f"{base_url}/v2/transcript/{transcript_id}" while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": print(transcript["text"]) break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) ``` ```python {8} expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig audio_file = "https://assembly.ai/prompt-8" config = TranscriptionConfig( language_detection=True, prompt="League of Legends roles", ) transcriber = Transcriber(api_key="") transcript = transcriber.transcribe(audio_file, config) print(transcript.text) ``` ```javascript {10} expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const data = { audio_url: "https://assembly.ai/prompt-8", language_detection: true, prompt: "League of Legends roles", }; const url = `${baseUrl}/v2/transcript`; let res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ```javascript {13} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); const audioFile = "https://assembly.ai/prompt-8"; const params = { audio: audioFile, language_detection: true, prompt: "League of Legends roles", }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); }; run(); ``` ### Prompting guide Contextual prompts work at three levels of specificity. Use the least specific level that covers your use case, and add detail when your audio contains uncommon names or terms the model can't otherwise know. | Level | Length | What it contains | Example | | --- | --- | --- | --- | | **Domain** | 2–5 words | The domain only | `Medical consultation call.` | | **Scenario** | 5–15 words | What the conversation is about | `Cardiology consultation about chest pain symptoms.` | | **Detailed** | 20–50 words | Full description, including names, products, or identifiers | `Cardiology consultation between Dr. Smith and an elderly patient regarding recurring chest pain, ECG results, and medication adjustment for hypertension.` | Guidelines for writing contextual prompts: - Write plain, complete sentences that describe the audio - Keep it to one short block of text. Don't pack lists of keywords into the contextual prompt ## Keyterms prompting Keyterms prompting allows you to provide up to 1,000 words or phrases (maximum 6 words per phrase) using the `keyterms_prompt` parameter to improve transcription accuracy for those terms and related variations or contextually similar phrases. Here is an example showing how you can use keyterms prompting to improve transcription accuracy for a name with distinctive spelling and formatting. Without keyterms prompting: ```txt Hi, this is Kelly Byrne Donahue ``` With keyterms prompting: ```txt Hi, this is Kelly Byrne-Donoghue ``` ```python {10} expandable import requests import time base_url = "https://api.assemblyai.com" headers = {"authorization": ""} data = { "audio_url": "https://assemblyaiassets.com/audios/keyterms_prompting.wav", "language_detection": True, "keyterms_prompt": ["Kelly Byrne-Donoghue"] } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) if response.status_code != 200: print(f"Error: {response.status_code}, Response: {response.text}") response.raise_for_status() transcript_response = response.json() transcript_id = transcript_response["id"] polling_endpoint = f"{base_url}/v2/transcript/{transcript_id}" while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": print(transcript["text"]) break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) ``` ```javascript {11} expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const data = { audio_url: "https://assemblyaiassets.com/audios/keyterms_prompting.wav", language_detection: true, keyterms_prompt: ["Kelly Byrne-Donoghue"], }; const url = `${baseUrl}/v2/transcript`; let res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ```python {7} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig audio_file = "https://assemblyaiassets.com/audios/keyterms_prompting.wav" config = TranscriptionConfig( language_detection=True, keyterms_prompt=["Kelly Byrne-Donoghue"] ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) print(transcript.text) ``` ```javascript {12} import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); const audioFile = "https://assemblyaiassets.com/audios/keyterms_prompting.wav"; const params = { audio: audioFile, language_detection: true, keyterms_prompt: ["Kelly Byrne-Donoghue"], }; const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); ``` **Keyword count limits** While we support up to 1000 key words and phrases, actual capacity may be lower due to internal tokenization and implementation constraints. Key points to remember: - Each word in a multi-word phrase counts towards the 1000 keyword limit - Capitalization affects capacity (uppercase tokens consume more than lowercase) - Longer words consume more capacity than shorter words For optimal results, use shorter phrases when possible and be mindful of your total token count when approaching the keyword limit. ## Need help? If you'd like help building or optimizing a prompt for your audio, our team can help: open a live chat or email us via the widget in the bottom-right corner ([contact info](https://www.assemblyai.com/contact/support)). --- # Medical Mode URL: https://www.assemblyai.com/docs/pre-recorded-audio/medical-mode Source: docs/pre-recorded-audio/medical-mode.mdx Navigation: Pre-recorded STT > Features Description: Improve transcription accuracy for medical terminology in pre-recorded audio Medical Mode is an add-on that enhances transcription accuracy for medical terminology — including medication names, procedures, conditions, and dosages. It is optimized for medical entity recognition to correct terms that other models frequently get wrong. Medical Mode can be used with all of our Pre-recorded STT models. Enable Medical Mode by setting the `domain` parameter to `"medical-v1"`. No changes to your existing pipeline are required. ## Quickstart To enable Medical Mode, set `domain` to `"medical-v1"` in the POST request body: ```python {11} expandable import requests import time base_url = "https://api.assemblyai.com" headers = {"authorization": ""} data = { "audio_url": "https://assembly.ai/lispro", "language_detection": True, "domain": "medical-v1" } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) if response.status_code != 200: print(f"Error: {response.status_code}, Response: {response.text}") response.raise_for_status() transcript_response = response.json() transcript_id = transcript_response["id"] polling_endpoint = f"{base_url}/v2/transcript/{transcript_id}" while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": print(transcript["text"]) break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) ``` To enable Medical Mode, set `domain` to `"medical-v1"` in the transcription config. ```python {12} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # You can use a local filepath: # audio_file = "./example.mp3" # Or use a publicly-accessible URL: audio_file = "https://assembly.ai/lispro" config = TranscriptionConfig( language_detection=True, domain="medical-v1", ) transcriber = Transcriber(api_key="") transcript = transcriber.transcribe(audio_file, config) print(transcript.text) ``` To enable Medical Mode, set `domain` to `"medical-v1"` in the POST request body: ```javascript {10} expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const data = { audio_url: "https://assembly.ai/lispro", language_detection: true, domain: "medical-v1", }; const url = `${baseUrl}/v2/transcript`; let res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` To enable Medical Mode, set `domain` to `"medical-v1"` in the transcription config. ```javascript {17} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // You can use a local filepath: // const audioFile = "./example.mp3" // Or use a publicly-accessible URL: const audioFile = "https://assembly.ai/lispro"; const params = { audio: audioFile, language_detection: true, domain: "medical-v1", }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); }; run(); ``` ### Example output Without Medical Mode: ```plain I have here insulin to be used for both prandial mealtime and sliding scale is insulin lisprohumalog subcutaneously. ``` With Medical Mode, lisprohumalog is updated to Lispro (Humalog) - following the standard medical convention of writing the generic name first, with the brand name in parentheses. ```plain I have here insulin to be used for both prandial mealtime and sliding scale is insulin Lispro (Humalog) subcutaneously. ``` ## Use cases Medical Mode is designed for healthcare AI applications where accurate medical terminology is critical: - **Ambient clinical documentation** — Capture medication names, dosages, and clinical terms correctly in real-time scribing workflows. - **AI-powered clinical notes** — Generate clean transcripts for downstream LLMs producing SOAP notes, discharge summaries, and referral letters. - **Front-office automation** — Handle drug names, provider names, and clinic-specific terminology in scheduling calls, insurance verification, and voice agents. - **Multi-speaker clinical conversations** — Combine with [Speaker Diarization](/pre-recorded-audio/label-speakers) for provider/patient separation in telehealth, therapy documentation, and clinical settings. ## Combine with other features Medical Mode works alongside other transcription features. You can combine it with: - [Speaker Diarization](/pre-recorded-audio/label-speakers) to identify who said what in clinical conversations - [Keyterms Prompting](/pre-recorded-audio/universal-3-5-pro/prompting#keyterms-prompting) to further boost accuracy for specific medical terms unique to your use case - [PII Redaction](/guardrails/redact-pii-from-transcripts) to redact sensitive patient information from transcripts ```python data = { "audio_url": "", "language_detection": True, "domain": "medical-v1", "speaker_labels": True, "keyterms_prompt": ["Lisinopril", "Metformin", "Humalog"] } ``` ```python config = TranscriptionConfig( language_detection=True, domain="medical-v1", speaker_labels=True, keyterms_prompt=["Lisinopril", "Metformin", "Humalog"], ) ``` ```javascript const data = { audio_url: "", language_detection: true, domain: "medical-v1", speaker_labels: true, keyterms_prompt: ["Lisinopril", "Metformin", "Humalog"], }; ``` ```javascript const params = { audio: audioFile, language_detection: true, domain: "medical-v1", speaker_labels: true, keyterms_prompt: ["Lisinopril", "Metformin", "Humalog"], }; ``` Medical Mode supports English, Spanish, German, and French. If you use Medical Mode with an unsupported language, the API ignores the `domain` parameter and returns a warning indicating that Medical Mode was not applied: `"Skipped medical-v1 domain correction because the language is not supported"` Your transcript is still returned using standard transcription, and you will not be charged for Medical Mode. ## HIPAA compliance AssemblyAI offers a Business Associate Agreement (BAA) for customers who need to process Protected Health Information (PHI). AssemblyAI is SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0 certified. Medical Mode does not change existing data handling or retention policies. Paid customers can review and sign our standard online BAA self-serve, at no additional cost, from the [**Data Controls** page](https://www.assemblyai.com/dashboard/settings/data-controls) in the dashboard. You must have an upgraded (paid) account to see the BAA option. See [Data Controls](/data-controls#business-associate-agreement-baa) for details, or [contact sales](https://www.assemblyai.com/contact-sales) for enterprise pricing. --- # Speaker Diarization URL: https://www.assemblyai.com/docs/pre-recorded-audio/label-speakers Source: docs/pre-recorded-audio/label-speakers.mdx Navigation: Pre-recorded STT > Features Description: Add speaker labels to your transcript ## Overview Speaker diarization identifies individual speakers in your audio and labels each segment of the transcript with the speaker. When enabled, the transcript is returned as a list of utterances, where each utterance represents an uninterrupted segment of speech from a single speaker. Accuracy improves the more each speaker talks as the model accumulates embedding context. For best results, each speaker should have at least 30 seconds of continuous speech. ### Quickstart To enable Speaker Diarization, set `speaker_labels` to `True` in the POST request body: ```python {21} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "speaker_labels": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) for utterance in transcription_result['utterances']: print(f"Speaker {utterance['speaker']}: {utterance['text']}") ``` To enable Speaker Diarization, set `speaker_labels` to `True` in the transcription config. ```python {14} expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # You can use a local filepath: # audio_file = "./example.mp3" # Or use a publicly-accessible URL: audio_file = ( "https://assembly.ai/wildfires.mp3" ) config = TranscriptionConfig( language_detection=True, speaker_labels=True, ) transcriber = Transcriber(api_key="") transcript = transcriber.transcribe(audio_file, config) for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` To enable Speaker Diarization, set `speaker_labels` to `true` in the POST request body: ```javascript {26} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./audio/audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, speaker_labels: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { for (const utterance of transcriptionResult.utterances) { console.log(`Speaker ${utterance.speaker}: ${utterance.text}`); } break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` To enable Speaker Diarization, set `speaker_labels` to `true` in the transcription config. ```javascript {17} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // You can use a local filepath: // const audioFile = "./example.mp3" // Or use a publicly-accessible URL: const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, speaker_labels: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); for (const utterance of transcript.utterances ?? []) { console.log(`Speaker ${utterance.speaker}: ${utterance.text}`); } }; run(); ``` ### Configuration You can constrain the number of speakers with `speakers_expected`, or with `min_speakers_expected`/`max_speakers_expected`. These are hard boundaries on the number of speaker labels, not hints: `max_speakers_expected` is a strict cap — if more people speak than that, the additional speakers are merged into existing labels — and `min_speakers_expected` is a strict floor. - If you know the **exact** number of speakers, set `speakers_expected`. - If you have a rough idea, use `min_speakers_expected` and `max_speakers_expected` to set a range. Set `max_speakers_expected` a little higher than the number of speakers you expect so the model has room to identify any additional speakers. Setting it too high can cause the model to over-split and return more speaker labels than are actually present.
Key Type Default Description
`speaker_labels` boolean `false` Enable Speaker Diarization.
`speakers_expected` number Set the exact number of speakers.
`speaker_options` object Set range of possible speakers.
`speaker_options.min_speakers_expected` number A hard lower limit on the number of speaker labels. The model won't return fewer speakers than this.
`speaker_options.max_speakers_expected` number **0–2 minutes**: —
**2–10 minutes**: 10 speakers
**10+ minutes**: 30 speakers
A hard upper limit on the number of speaker labels. If more people speak than this, the additional speakers are merged into existing labels. Give the model a little headroom above the number of speakers you expect; setting it too high can cause over-splitting and return more speakers than are actually present.
Only set `speakers_expected` when you are certain of the exact speaker count. If you're unsure, use `min_speakers_expected` and `max_speakers_expected` to describe a range instead — providing an incorrect exact count can negatively affect diarization accuracy. ### Reading the response When diarization is enabled, the transcript includes an `utterances` array in place of a single text block. Each object in the array represents one uninterrupted segment of speech from a single speaker. The `utterances` array contains objects with the following fields: | Field | Type | Description | | ------------- | ------ | ------------------------------------------------------------------------------------------------------ | | `speaker` | string | The speaker label, assigned as sequential letters such as A, B, and C. | | `text` | string | The transcribed text for the utterance. | | `start` | number | The start time of the utterance in milliseconds. | | `end` | number | The end time of the utterance in milliseconds. | | `confidence` | number | A score between 0 and 1 indicating the model's confidence in the transcribed text. | | `words` | array | A word-level breakdown of the utterance. | Each object in the `words` array contains the following fields: | Field | Type | Description | | ------------ | ------ | ----------------------------------------------------------------------------------- | | `text` | string | The transcribed word. | | `speaker` | string | The speaker label for the word. | | `start` | number | The start time of the word in milliseconds. | | `end` | number | The end time of the word in milliseconds. | | `confidence` | number | A score between 0 and 1 indicating the model's confidence in the word. | ### Identify speakers by name Speaker Diarization assigns generic labels like "Speaker A" and "Speaker B" to each speaker. If you want to replace these labels with actual names or roles, you can use Speaker Identification to transform your transcript. **Before Speaker Identification:** ```txt Speaker A: Good morning, and welcome to the show. Speaker B: Thanks for having me. ``` **After Speaker Identification:** ```txt Michel Martin: Good morning, and welcome to the show. Peter DeCarlo: Thanks for having me. ``` The following example shows how to transcribe audio with Speaker Diarization and then apply Speaker Identification to replace the generic speaker labels with actual names. ```python maxLines=30 expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } audio_url = "https://assembly.ai/wildfires.mp3" # Configure transcript with speaker diarization and speaker identification data = { "audio_url": audio_url, "language_detection": True, "speaker_labels": True, "speech_understanding": { "request": { "speaker_identification": { "speaker_type": "name", "known_values": ["Michel Martin", "Peter DeCarlo"] } } } } # Submit the transcription request response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) transcript_id = response.json()["id"] polling_endpoint = base_url + f"/v2/transcript/{transcript_id}" # Poll for transcription results while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) # Print utterances with identified speaker names for utterance in transcript["utterances"]: print(f"{utterance['speaker']}: {utterance['text']}") ``` ```javascript maxLines=30 expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", "content-type": "application/json", }; const audioUrl = "https://assembly.ai/wildfires.mp3"; // Configure transcript with speaker diarization and speaker identification const data = { audio_url: audioUrl, language_detection: true, speaker_labels: true, speech_understanding: { request: { speaker_identification: { speaker_type: "name", known_values: ["Michel Martin", "Peter DeCarlo"], }, }, }, }; async function main() { // Submit the transcription request const response = await fetch(`${baseUrl}/v2/transcript`, { method: "POST", headers: headers, body: JSON.stringify(data), }); if (!response.ok) throw new Error(`Error: ${response.status}`); const { id: transcriptId } = await response.json(); const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; // Poll for transcription results while (true) { const pollingResponse = await fetch(pollingEndpoint, { headers }); if (!pollingResponse.ok) throw new Error(`Error: ${pollingResponse.status}`); const transcript = await pollingResponse.json(); if (transcript.status === "completed") { // Print utterances with identified speaker names for (const utterance of transcript.utterances) { console.log(`${utterance.speaker}: ${utterance.text}`); } break; } else if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } } main().catch(console.error); ``` For more details on Speaker Identification, including how to identify speakers by role and how to apply it to existing transcripts, see the [Speaker Identification guide](/speech-understanding/speaker-identification). ### Best practices for accurate diarization Follow these tips to get the best results from Speaker Diarization: - **Ensure sufficient speech per speaker.** Each speaker should speak for at least 30 seconds uninterrupted. The model may struggle to create separate clusters for speakers who only contribute short phrases like "Yeah", "Right", or "Sounds good". - **Minimize cross-talk.** Overlapping speech between speakers can reduce diarization accuracy. Where possible, ensure speakers take turns. - **Reduce background noise.** Background noise, echoes, or playback of recorded audio during a conversation can interfere with speaker separation. - **Use `speaker_options` instead of `speakers_expected` when uncertain.** Only use `speakers_expected` when you are confident about the exact number of speakers. If this number is incorrect, the model may produce random splits of single-speaker segments or merge multiple speakers into one. It's generally recommended to use `min_speakers_expected` and set `max_speakers_expected` slightly higher (e.g., `min_speakers_expected` + 2) to allow flexibility. - **Avoid setting `max_speakers_expected` too high.** Setting the maximum too high may reduce accuracy, causing sentences from the same speaker to be split across multiple speaker labels. - **Be aware of speaker similarity.** If speakers sound similar, the model may have difficulty distinguishing between them. --- # Multichannel Transcription URL: https://www.assemblyai.com/docs/pre-recorded-audio/transcribe-multiple-audio-channels Source: docs/pre-recorded-audio/transcribe-multiple-audio-channels.mdx Navigation: Pre-recorded STT > Features Description: Multichannel Transcription documentation. If you have a multichannel audio file with multiple speakers, you can transcribe each of them separately. The response includes an `audio_channels` property with the number of different channels, and an additional `utterances` property, containing a list of turn-by-turn utterances. Each utterance contains channel information, starting at 1. Additionally, each word in the `words` array contains the channel identifier. ## Quickstart ```python title="Python SDK" {8} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True, multichannel=True ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") for utterance in transcript.utterances: print(f"Channel {utterance.speaker}: {utterance.text}") ``` ```python title="Python" {20} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "multichannel": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) for utterance in transcription_result['utterances']: print(f"Channel {utterance['speaker']}: {utterance['text']}") ``` ```javascript title="JavaScript SDK" {13} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, multichannel: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); for (const utterance of transcript.utterances ?? []) { console.log(`Channel ${utterance.speaker}: ${utterance.text}`); } }; run(); ``` ```javascript title="JavaScript" {21} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, multichannel: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { for (const utterance of transcriptionResult.utterances) { console.log(`Channel ${utterance.speaker}: ${utterance.text}`); } break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` Multichannel audio increases the transcription time by approximately 40%. ## Per-channel diarization If you have a multichannel audio file where individual channels may contain multiple speakers, you can combine `multichannel` and `speaker_labels` to perform diarization within each channel. When using `multichannel` with `speaker_labels`, the `speaker_options` parameters (`min_speakers_expected` and `max_speakers_expected`) are applied **per channel**, not globally across the entire file. For example, setting `min_speakers_expected: 5` and `max_speakers_expected: 7` on a 5-channel file means the model will find 5–7 speakers on _each_ channel, resulting in 25–35 total speakers. Adjust your speaker options accordingly when using multichannel transcription. When both parameters are enabled: - Channels are labeled numerically (1, 2, 3, etc.) - Speakers within each channel are labeled alphabetically (A, B, C, etc.) - The combined speaker label format is `{channel}{speaker}` (e.g., "1A", "1B", "2A") For example, if channel 1 has two speakers and channel 2 has one speaker, the labels would be: - First speaker on channel 1: `1A` - Second speaker on channel 1: `1B` - First speaker on channel 2: `2A` ```python title="Python SDK" {8-9} expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True, multichannel=True, speaker_labels=True ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` ```python title="Python" {20-21} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "multichannel": True, "speaker_labels": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) for utterance in transcription_result['utterances']: print(f"Speaker {utterance['speaker']}: {utterance['text']}") ``` ```javascript title="JavaScript SDK" {13-14} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, multichannel: true, speaker_labels: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); for (const utterance of transcript.utterances ?? []) { console.log(`Speaker ${utterance.speaker}: ${utterance.text}`); } }; run(); ``` ```javascript title="JavaScript" {21-22} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, multichannel: true, speaker_labels: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { for (const utterance of transcriptionResult.utterances) { console.log(`Speaker ${utterance.speaker}: ${utterance.text}`); } break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` --- # Code Switching URL: https://www.assemblyai.com/docs/pre-recorded-audio/code-switching Source: docs/pre-recorded-audio/code-switching.mdx Navigation: Pre-recorded STT > Features Description: Code Switching documentation. Transcribe audio containing multiple languages with code switching detection. This feature enables accurate transcription of conversations where speakers naturally switch between languages during conversations. ## Universal-3.5 Pro (Recommended) Universal-3.5 Pro natively handles code switching across 18 languages. Set `language_detection` as `True` and the model follows speakers as they shift mid-sentence between languages like English and Spanish, French, Hindi or Mandarin, preserving exactly what was said without translating everything into a single language. See Universal-3.5 Pro code switching in action. ### **English \<\> French** ```txt I said something like, j'ai dit à mes étudiants que it's time to really pay attention to what the idea of code-switching is. ``` ### **English \<\> Hindi** ```txt मेरा रुकने का तो बहुत मन है, but I have an exam to give tomorrow. ``` ### **English \<\> Mandarin** ```txt But this sentence, 我父母不工作了, you can see the 了 at the end of the sentence indicates that the situation now, my parents don't work, is different from what it was before. ``` ### Quickstart ```python title="Python SDK" {7} expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local-file.mp3" audio_file = "https://assembly.ai/code-switching-3" config = TranscriptionConfig( language_detection=True, ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" {10} expandable import requests import time base_url = "https://api.assemblyai.com" headers = {"authorization": ""} data = { "audio_url": "https://assembly.ai/code-switching-3", "language_detection": True, } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) if response.status_code != 200: print(f"Error: {response.status_code}, Response: {response.text}") response.raise_for_status() transcript_response = response.json() transcript_id = transcript_response["id"] polling_endpoint = f"{base_url}/v2/transcript/{transcript_id}" while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": print(transcript["text"]) break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" {14} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // You can use a local filepath: // const audioFile = "./local-file.mp3"; // Or use a publicly-accessible URL: const audioFile = "https://assembly.ai/code-switching-3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { console.error(`Transcription failed: ${transcript.error}`); process.exit(1); } console.log(`\nFull Transcript:\n\n${transcript.text}\n`); }; run(); ``` ```javascript title="JavaScript" {10} expandable const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const data = { audio_url: "https://assembly.ai/code-switching-3", language_detection: true, }; const url = `${baseUrl}/v2/transcript`; let res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ## Universal-2 While Universal-2 supports code switching, we recommend upgrading to Universal-3.5 Pro for best results. To enable code switching on Universal-2, set `speech_models` to `universal-2`, `language_detection` to `true`, and `code_switching` to `true` inside the `language_detection_options` parameter. ### Quickstart ```python title="Python SDK" {11} expandable from assemblyai import LanguageDetectionOptions from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig audio_file = "./bilingual-audio.mp3" # audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( speech_models=["universal-2"], language_detection=True, language_detection_options=LanguageDetectionOptions( code_switching=True ) ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" {22} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./bilingual-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, "speech_models": ["universal-2"], "language_detection": True, "language_detection_options": { "code_switching": True }, } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" {18} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // You can use a local filepath: const audioFile = "./bilingual-audio.mp3"; // Or use a publicly-accessible URL: // const audioFile = ""; const params = { audio: audioFile, speech_models: ["universal-2"], language_detection: true, language_detection_options: { code_switching: true, }, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { console.error(`Transcription failed: ${transcript.error}`); process.exit(1); } console.log(`\nFull Transcript:\n\n${transcript.text}\n`); }; run(); ``` ```javascript title="JavaScript" {22} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./bilingual-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, speech_models: ["universal-2"], language_detection: true, language_detection_options: { code_switching: true, }, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ### Example API Response When enabling code switching with automatic language detection, the two detected language codes with the highest confidence and their confidence will be included in the transcript JSON. ```json "language_detection_results": { "code_switching_languages": [ {"language": "en", "confidence": 0.8}, {"language": "es", "confidence": 0.7} ] } ``` ### Manually Setting Language Codes To manually set the [language codes](/pre-recorded-audio/code-switching), you can use the `language_codes` parameter. A max of two language codes can be set and one code must be `"en"`. For example, if your file contains both English and Spanish, it would be `"language_codes": ["en", "es"]`. ```python title="Python SDK" {8} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig audio_file = "./bilingual-audio.mp3" # audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( speech_models=["universal-2"], language_codes=["en", "es"] # English-Spanish code switching ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" {20} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./bilingual-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, "speech_models": ["universal-2"], "language_codes": ["en", "es"] # English-Spanish code switching } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" {16} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // You can use a local filepath: const audioFile = "./bilingual-audio.mp3"; // Or use a publicly-accessible URL: // const audioFile = ""; const params = { audio: audioFile, speech_models: ["universal-2"], language_codes: ["en", "es"], // English-Spanish code switching }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { console.error(`Transcription failed: ${transcript.error}`); process.exit(1); } console.log(`\nFull Transcript:\n\n${transcript.text}\n`); }; run(); ``` ```javascript title="JavaScript" {20} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./bilingual-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, speech_models: ["universal-2"], language_codes: ["en", "es"], // English-Spanish code switching }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ### Code Switching Confidence Threshold The `code_switching_confidence_threshold` parameter controls how the model routes transcription when multiple languages are detected. When code switching is enabled, the model detects up to two languages per audio file and assigns each a confidence score. **Code Switching Routing Behavior** This parameter controls routing, not rejection. Audio is always transcribed, even if confidence scores do not meet this threshold. To return an error instead of a low-confidence transcription, you can use [language_confidence_threshold](/pre-recorded-audio/language-detection#set-a-language-confidence-threshold) alongside this parameter. The threshold determines which language is used for transcription using the following logic: - If the non-English language's confidence score **meets or exceeds** the threshold, the audio is routed to that non-English language model. - If the non-English language's confidence score **falls below** the threshold, the audio is routed to whichever language has the **highest overall confidence**, which may be English or non-English. - If both detected languages are non-English, the audio is always routed to whichever has the higher confidence score, regardless of the threshold. **Code Switching Default** By default, the `code_switching_confidence_threshold` parameter is set to `0.3`. If you would like to disable this, make sure to set this parameter to `0`. Setting `code_switching_confidence_threshold` to `0` means the non-English language is always used for routing, even with very low confidence. ```python title="Python SDK" {12} expandable from assemblyai import LanguageDetectionOptions from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig audio_file = "./bilingual-audio.mp3" # audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( speech_models=["universal-2"], language_detection=True, language_detection_options=LanguageDetectionOptions( code_switching=True, code_switching_confidence_threshold=0.5 # Optional parameter - this is set to 0.3 by default ) ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" {23} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./bilingual-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, "speech_models": ["universal-2"], "language_detection": True, "language_detection_options": { "code_switching": True, "code_switching_confidence_threshold": 0.5 # Optional parameter - this is set to 0.3 by default }, } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" {19} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // You can use a local filepath: const audioFile = "./bilingual-audio.mp3"; // Or use a publicly-accessible URL: // const audioFile = ""; const params = { audio: audioFile, speech_models: ["universal-2"], language_detection: true, language_detection_options: { code_switching: true, code_switching_confidence_threshold: 0.5, // Optional parameter - this is set to 0.3 by default }, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { console.error(`Transcription failed: ${transcript.error}`); process.exit(1); } console.log(`\nFull Transcript:\n\n${transcript.text}\n`); }; run(); ``` ```javascript title="JavaScript" {23} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./bilingual-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, speech_models: ["universal-2"], language_detection: true, language_detection_options: { code_switching: true, code_switching_confidence_threshold: 0.5, // Optional parameter - this is set to 0.3 by default }, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ## Support for 99 languages Universal-3.5 Pro supports 18 languages, and for anything outside that set, the system automatically falls back to Universal-2, giving you coverage across 99 languages total without any extra configuration. | Model | Supported languages | | --- | --- | | `universal-3-5-pro` | Global English, Australian English, British English, US English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Mandarin, Norwegian, Swedish, Turkish, Vietnamese | | `universal-2` | Global English, Australian English, British English, US English, Spanish, French, German, Italian, Portuguese, Dutch, Hindi, Japanese, Chinese, Finnish, Korean, Polish, Russian, Turkish, Ukrainian, Vietnamese, Afrikaans, Albanian, Amharic, Arabic, Armenian, Assamese, Azerbaijani, Bashkir, Basque, Belarusian, Bengali, Bosnian, Breton, Bulgarian, Burmese, Catalan, Croatian, Czech, Danish, Estonian, Faroese, Galician, Georgian, Greek, Gujarati, Haitian, Hausa, Hawaiian, Hebrew, Hungarian, Icelandic, Indonesian, Javanese, Kannada, Kazakh, Khmer, Lao, Latin, Latvian, Lingala, Lithuanian, Luxembourgish, Macedonian, Malagasy, Malay, Malayalam, Maltese, Maori, Marathi, Mongolian, Nepali, Norwegian, Norwegian Nynorsk, Occitan, Panjabi, Pashto, Persian, Romanian, Sanskrit, Serbian, Shona, Sindhi, Sinhala, Slovak, Slovenian, Somali, Sundanese, Swahili, Swedish, Swiss German, Tagalog, Tajik, Tamil, Tatar, Telugu, Thai, Tibetan, Turkmen, Urdu, Uzbek, Welsh | --- # Automatic Language Detection URL: https://www.assemblyai.com/docs/pre-recorded-audio/language-detection Source: docs/pre-recorded-audio/language-detection.mdx Navigation: Pre-recorded STT > Features > Languages Description: Detect the dominant language and automatically route your request to the best available speech model for that language. Automatic language detection identifies the dominant language in your audio and routes the request to the best available model based on the detected language and the models you specify in `speech_models`. You can check which model processed your request using the `speech_model_used` field in the response. For best results, your audio should contain at least 15–90 seconds of spoken audio. ## Quickstart ```python title="Python SDK" for="python-sdk" highlight={13} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) print(transcript.text) print(transcript.json_response["language_code"]) ``` ```python title="Python" for="python" highlight={20} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") print(f"Language Code: {transcription_result['language_code']}") print(f"Text: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" for="javascript-sdk" highlight={13} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); console.log(transcript.language_code); }; run(); ``` ```javascript title="JavaScript" for="javascript" highlight={21} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); console.log(transcriptionResult.language_code); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ## Set a list of expected languages If you're confident the audio is in one of a few languages, provide that list via `language_detection_options.expected_languages`. Detection is restricted to these candidates and the model will choose the language with the highest confidence from this list. This can eliminate scenarios where Automatic Language Detection selects an unexpected language for transcription. - Use our [language codes](/pre-recorded-audio/supported-languages) (e.g., `"en"`, `"es"`, `"fr"`). - If `expected_languages` is not specified, it is set to `["all"]` by default. ```python title="Python SDK" for="python-sdk" highlight={13} expandable from assemblyai import LanguageDetectionOptions from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" options = LanguageDetectionOptions( expected_languages=["en", "es", "fr", "de"], fallback_language="auto" ) config = TranscriptionConfig( language_detection=True, language_detection_options=options ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) print(transcript.text) print(transcript.json_response["language_code"]) ``` ```python title="Python" for="python" highlight={22} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "language_detection_options": { "expected_languages": ["en", "es", "fr", "de"], "fallback_language": "auto" } } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") print(f"Language Code: {transcription_result['language_code']}") print(f"Text: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" for="javascript-sdk" highlight={15} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, language_detection_options: { expected_languages: ["en", "es", "fr", "de"], fallback_language: "auto", }, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); console.log(transcript.language_code); }; run(); ``` ```javascript title="JavaScript" for="javascript" highlight={22} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, language_detection_options: { expected_languages: ["en", "es", "fr", "de"], fallback_language: "auto", }, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); console.log(transcriptionResult.language_code); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ## Choose a fallback language Control what language transcription should fall back to when detection cannot confidently select a language from the `expected_languages` list. - Set `language_detection_options.fallback_language` to a specific language code (e.g., `"en"`). - `fallback_language` must be one of the language codes in `expected_languages` or `"auto"`. - When `fallback_language` is unspecified, it is set to `"auto"` by default. This tells our model to choose the fallback language from `expected_languages` with the highest confidence score. ```python title="Python SDK" for="python-sdk" highlight={14} expandable from assemblyai import LanguageDetectionOptions from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" options = LanguageDetectionOptions( expected_languages=["en", "es", "fr", "de"], fallback_language="auto" ) config = TranscriptionConfig( language_detection=True, language_detection_options=options ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) print(transcript.text) print(transcript.json_response["language_code"]) ``` ```python title="Python" for="python" highlight={22} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "language_detection_options": { "expected_languages": ["en", "es", "fr", "de"], "fallback_language": "auto" } } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") print(f"Language Code: {transcription_result['language_code']}") print(f"Text: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" for="javascript-sdk" highlight={15} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, language_detection_options: { expected_languages: ["en", "es", "fr", "de"], fallback_language: "auto", }, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); console.log(transcript.language_code); }; run(); ``` ```javascript title="JavaScript" for="javascript" highlight={22} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, language_detection_options: { expected_languages: ["en", "es", "fr", "de"], fallback_language: "auto", }, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); console.log(transcriptionResult.language_code); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ## Choose a locale When English is detected by automatic language detection, the transcript is rendered in the default US English spelling. Set your preferred English locale to render it in a regional spelling variant instead. - Set `language_detection_options.localization` to a locale code. The supported values are `en_au` (Australian English) and `en_uk` (British English). - When English is detected, the transcript uses that locale's spelling — for example `colour`, `behaviour`, and `organise`. - `language_code` in the response returns the region-aware code (such as `en_au`) instead of the base `en`. ```python title="Python" for="python" highlight={21} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "language_detection_options": { "localization": ["en_au"] } } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(transcription_result['text']) print(transcription_result['language_code']) break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript" for="javascript" highlight={24} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, language_detection_options: { localization: ["en_au"], }, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); console.log(transcriptionResult.language_code); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` **Things to know** - **Only applies to English.** If the detected language isn't English, transcription proceeds as normal in the detected language. The locale is applied only when the detected language is English; otherwise the request is unaffected. - **English only.** `en_au` and `en_uk` are the only supported locales today. Base and default-region codes such as `en` and `en_us` aren't localization variants and return a `400`. - **One locale per base language.** Passing both `en_au` and `en_uk` returns a `400`. - **Requires automatic language detection.** `localization` only applies when `language_detection` is `true`. If you already know the language, set `language_code` to `en_au` or `en_uk` directly instead. - **Spelling only.** Localization affects spelling, not vocabulary or grammar. ## Confidence score If language detection is enabled, the API returns a confidence score for the detected language. The score ranges from 0.0 (low confidence) to 1.0 (high confidence). ```python title="Python SDK" for="python-sdk" highlight={7,13} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) print(transcript.text) print(transcript.json_response["language_confidence"]) ``` ```python title="Python" for="python" highlight={33} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") print(f"Language Confidence: {transcription_result['language_confidence']}") print(f"Text: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" for="javascript-sdk" highlight={19} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); console.log(transcript.language_confidence); }; run(); ``` ```javascript title="JavaScript" for="javascript" highlight={36} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); console.log(transcriptionResult.language_confidence); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ## Set a language confidence threshold You can set the confidence threshold that must be reached if language detection is enabled. An error will be returned if the language confidence is below this threshold. Valid values are in the range [0,1] inclusive. ```python title="Python SDK" for="python-sdk" highlight={7,13} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True, language_confidence_threshold=0.8 ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") else: print(transcript.json_response["language_confidence"]) print(transcript.text) ``` ```python title="Python" for="python" highlight={20} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "language_confidence_threshold": 0.8 } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") print(f"Text: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" for="javascript-sdk" highlight={13} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, language_confidence_threshold: 0.8, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } console.log(transcript.text); console.log(transcript.language_confidence); }; run(); ``` ```javascript title="JavaScript" for="javascript" highlight={21} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, language_confidence_threshold: 0.8, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); console.log(transcriptionResult.language_confidence); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` If the `language_confidence_threshold` you specify is not met you will receive an error message like `detected language 'bg', confidence 0.2949, is below the requested confidence threshold value of '0.4'`. ## Troubleshooting ### Accented speech detected as the wrong language Automatic Language Detection uses Whisper-based language identification, which can sometimes misidentify heavily accented speech as a different language. For example, English spoken with a strong accent may be detected as Finnish, Latvian, Latin, or Arabic. When this happens, the model might not just return a wrong language label -- it might also **transcribe the audio in the incorrectly detected language**. This effectively translates the speech rather than transcribing it, producing output in a language the speaker wasn't using. The exact transcription behavior can vary depending on the detected language and speech model used. ### Recommended mitigations **Use `expected_languages` to constrain detection (most effective).** If you know which languages your audio may contain, set `expected_languages` to only those languages. This prevents the model from selecting an unexpected language entirely. For example, if your application processes interviews in English, Spanish, and French: ```json { "language_detection": true, "language_detection_options": { "expected_languages": ["en", "es", "fr"], "fallback_language": "en" } } ``` Setting `fallback_language` to your most common language (e.g., `"en"`) ensures that if the model can't confidently choose between the expected languages, it defaults to the language most likely to produce a useful transcript. **Use `language_confidence_threshold` to reject low-confidence detections.** Setting a threshold (e.g., `0.7`) causes the API to return an error instead of a transcript when confidence is low. This helps catch some misdetections, but not cases where the model is confidently wrong. **Monitor `language_confidence` in responses.** Log the `language_code` and `language_confidence` fields from your transcript responses. Unexpected language codes or unusual confidence patterns can help you identify misdetection issues early and decide whether to retry with `expected_languages` or flag the transcript for review. --- # Set Language Manually URL: https://www.assemblyai.com/docs/pre-recorded-audio/set-language-manually Source: docs/pre-recorded-audio/set-language-manually.mdx Navigation: Pre-recorded STT > Features > Languages Description: Specify a dominant language via language_code to route your request to the best available speech model for that language. If you already know the dominant language of your audio file, you can use the `language_code` parameter to specify the language rather than using [Automatic language detection](/pre-recorded-audio/language-detection). ## Quickstart ```python title="Python SDK" highlight={7} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_code="es" ) transcript = Transcriber(api_key="", config=config).transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" highlight={19} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_code": "es" } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" highlight={12} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_code: "es", }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); }; run(); ``` ```javascript title="JavaScript" highlight={20} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_code: "es", }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` See the [Supported languages](/pre-recorded-audio/supported-languages) page for all supported languages and their codes. --- # Supported Languages URL: https://www.assemblyai.com/docs/pre-recorded-audio/supported-languages Source: docs/pre-recorded-audio/supported-languages.mdx Navigation: Pre-recorded STT > Features > Languages Description: Supported languages for pre-recorded audio speech-to-text models. AssemblyAI supports a wide range of languages across our speech-to-text models for pre-recorded audio. The available languages vary by model. Check out the [Models](/getting-started/models) page to learn more about our different models and how to choose the best one for your use case. See our [Model selection](/pre-recorded-audio/select-the-speech-model) page for more details on specifying a model in your request. ## Universal-3.5 Pro Universal-3.5 Pro supports the following 18 languages. To automatically fall back to Universal-2 for anything outside this set, set `speech_models` to `["universal-3-5-pro", "universal-2"]` and set `language_detection` to `true`. ### Regional dialects and variants Universal-3.5 Pro goes beyond standard language support with deep understanding of regional dialects and local variants. Whether your audio features Quebecois French, Mexican Spanish, or Brazilian Portuguese, the model accurately captures speech as it's naturally spoken — including colloquial expressions, local vocabulary, and accent-specific pronunciation patterns. **Dialect support** You do not need to specify a dialect code to get accurate dialect transcription. Universal-3.5 Pro automatically recognizes regional speech patterns when using the base language code (e.g., `fr` for all French dialects, `es` for all Spanish dialects). | Dialect / Variant | Description | | --- | --- | | American English | Standard US English, including regional variants (Southern, Midwestern, Northeastern) | | British English | UK English, including Received Pronunciation and regional accents | | Australian English | Australian English with local expressions and pronunciation | | Dialect / Variant | Description | | --- | --- | | Castilian Spanish | Standard Peninsular Spanish as spoken in central and northern Spain | | Mexican Spanish | Mexican Spanish with local vocabulary and pronunciation | | Argentine Spanish | Rioplatense Spanish with distinctive *voseo* and pronunciation | | Colombian Spanish | Colombian Spanish with regional speech patterns | | Chilean Spanish | Chilean Spanish with rapid speech patterns and local slang | | Caribbean Spanish | Cuban, Dominican, and Puerto Rican Spanish dialects | | Spanglish | English-Spanish code-mixing common in US bilingual communities | | Dialect / Variant | Description | | --- | --- | | Metropolitan French | Standard Parisian French | | Canadian French (Quebecois) | Quebec French with distinctive vocabulary, pronunciation, and expressions | | Belgian French | Belgian French with local vocabulary and pronunciation | | Dialect / Variant | Description | | --- | --- | | Brazilian Portuguese | Brazilian Portuguese with local vocabulary, pronunciation, and expressions | | European Portuguese | Standard Lisbon Portuguese with Iberian pronunciation | | Dialect / Variant | Description | | --- | --- | | Standard Italian | Standard Italian based on Tuscan-influenced speech | ## Universal-2 Universal-2 supports 99 languages. Pass the corresponding `language_code` in your transcription request to specify the language. ### Accuracy metrics The following groups Universal-2 languages by transcription accuracy, measured by Word Error Rate (WER). English, Spanish, French, German, Indonesian, Italian, Japanese, Dutch, Polish, Portuguese, Russian, Swedish, Turkish, Ukrainian, Catalan 10% to ≤25% WER)"> Arabic, Azerbaijani, Bulgarian, Bosnian, Mandarin Chinese, Czech, Danish, Greek, Estonian, Finnish, Galician, Hebrew, Hindi, Croatian, Hungarian, Korean, Macedonian, Malay, Norwegian, Romanian, Slovak, Swiss German, Tagalog, Thai, Urdu, Vietnamese 25% to ≤50% WER)"> Afrikaans, Belarusian, Welsh, Persian (Farsi), Armenian, Icelandic, Kazakh, Lithuanian, Latvian, Maori, Marathi, Slovenian, Swahili, Tamil 50% WER)"> Amharic, Assamese, Bengali, Gujarati, Hausa, Javanese, Georgian, Khmer, Kannada, Luxembourgish, Lingala, Lao, Malayalam, Mongolian, Maltese, Burmese, Nepali, Occitan, Punjabi, Pashto, Sindhi, Shona, Somali, Serbian, Telugu, Tajik, Uzbek, Yoruba ## Unsupported feature behavior Not all features are available for every language. If you enable a feature that isn't supported for the language of your audio, the API's behavior depends on how the language was specified: - **With `language_code` (manual):** The API rejects the request and returns an error, such as `"The following models are not available in this language: speaker_labels"`. This lets you catch configuration issues before processing. - **With `language_detection` (automatic):** The request completes normally, but any features that aren't supported for the detected language are silently omitted from the response. The transcription itself still succeeds. This is because the language isn't known until after the request is submitted, so the API can't validate feature compatibility upfront. To avoid unexpected results when using Automatic Language Detection, use the `language_code` and `language_confidence` fields in the response to verify the detected language and handle cases where a feature may not have been applied. --- # Custom Spelling URL: https://www.assemblyai.com/docs/pre-recorded-audio/correct-spelling-of-terms Source: docs/pre-recorded-audio/correct-spelling-of-terms.mdx Navigation: Pre-recorded STT > Features Description: Custom Spelling documentation. Custom Spelling lets you customize how words are spelled or formatted in the transcript. To use Custom Spelling, include `custom_spelling` in your transcription parameters. The parameter should be a list of dictionaries, with each dictionary specifying a mapping from a word or phrase to a new spelling or format of a word. ```python {20} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "custom_spelling": [ { "from": ["Decarlo"], "to": "DeCarlo" }, { "from": ["SQL"], "to": "Sequel" } ] } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` To use Custom Spelling, pass a dictionary to `set_custom_spelling()` on the transcription config. Each key-value pair specifies a mapping from a word or phrase to a new spelling or format of a word. The key specifies the new spelling or format, and the corresponding value is the word or phrase you want to replace. ```python {9} expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) config.set_custom_spelling( { "Gettleman": ["gettleman"], "SQL": ["Sequel"], } ) transcriber = Transcriber(api_key="", config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` To use Custom Spelling, include `custom_spelling` in your transcription parameters. The parameter should be an array of objects, with each object specifying a mapping from a word or phrase to a new spelling or format of a word. ```javascript {23} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, custom_spelling: [ { from: ["Decarlo"], to: "DeCarlo", }, { from: ["SQL"], to: "Sequel", }, ], }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` To use Custom Spelling, include `custom_spelling` in your transcription parameters. The parameter should be an array of objects, with each object specifying a mapping from a word or phrase to a new spelling or format of a word. ```javascript {13} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, custom_spelling: [ { from: ["Decarlo"], to: "DeCarlo", }, { from: ["Sequel"], to: "SQL", }, ], }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); }; run(); ``` The value in the `to` key is case-sensitive, but the value in the `from` key isn't. Additionally, the `to` key must only contain one word, while the `from` key can contain multiple words. --- # Filler Words URL: https://www.assemblyai.com/docs/pre-recorded-audio/include-filler-words Source: docs/pre-recorded-audio/include-filler-words.mdx Navigation: Pre-recorded STT > Features Description: Filler Words documentation. The following filler words are removed by default: - "um" - "uh" - "hmm" - "mhm" - "uh-huh" - "ah" - "huh" - "hm" - "m" If you want to keep filler words in the transcript, you can set the `disfluencies` to `true` in the transcription config. ```python title="Python SDK" highlight={6} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True, disfluencies=True ) transcriber = Transcriber(api_key="", config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" highlight={19} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True, "disfluencies": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" highlight={12} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, disfluencies: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); }; run(); ``` ```javascript title="JavaScript" highlight={19} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, disfluencies: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` --- # Word Search URL: https://www.assemblyai.com/docs/pre-recorded-audio/search-for-words-in-transcript Source: docs/pre-recorded-audio/search-for-words-in-transcript.mdx Navigation: Pre-recorded STT > Features Description: Word Search documentation. You can search through a completed transcript for a specific set of keywords, which is useful for quickly finding relevant information. The parameter can be a list of words, numbers, or phrases up to five words. ```python title="Python SDK" expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcriber = Transcriber(api_key="", config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") # Set the words you want to search for words = ["foo", "bar", "foo bar", "42"] matches = transcript.word_search(words) for match in matches: print(f"Found '{match.text}' {match.count} times in the transcript") ``` ```python title="Python" expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) words = ['foo', 'bar', 'foo bar', '42'] word_search_endpoint = f"{base_url}/v2/transcript/{transcript_id}/word-search?words={','.join(words)}" response_json = requests.get(word_search_endpoint, headers=headers).json() for match in response_json['matches']: print(f"Found '{match['text']}' {match['count']} times in the transcript") ``` ```javascript title="JavaScript SDK" expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); // Set the words you want to search for. const words = ["foo", "bar", "foo bar", "42"]; const { matches } = await client.transcripts.wordSearch(transcript.id, words); for (const match of matches) { console.log(`Found '${match.text}' ${match.count} times in the transcript`); } }; run(); ``` ```javascript title="JavaScript" expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcript = await res.json(); const transcriptId = transcript.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } const words = ["foo", "bar", "foo bar", "42"]; res = await fetch(`${baseUrl}/v2/transcript/${transcriptId}/word-search?words=${words.join(",")}`, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); for (const match of response.matches) { console.log(`Found '${match.text}' ${match.count} times in the transcript`); } ``` --- # Set the Start and End of the Transcript URL: https://www.assemblyai.com/docs/pre-recorded-audio/set-the-start-and-end-of-the-transcript Source: docs/pre-recorded-audio/set-the-start-and-end-of-the-transcript.mdx Navigation: Pre-recorded STT > Features Description: Set the Start and End of the Transcript documentation. If you only want to transcribe a portion of your file, you can set the `audio_start_from` and the `audio_end_at` parameters in your transcription config. ```python title="Python SDK" highlight={7-8} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( audio_start_from=5000, # The start time of the transcription in milliseconds audio_end_at=15000 # The end time of the transcription in milliseconds ) transcriber = Transcriber(api_key="", config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" highlight={19,20} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "audio_start_from": 5000, # The start time of the transcription in milliseconds "audio_end_at": 15000 # The end time of the transcription in milliseconds } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" highlight={12-13} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, audio_start_from: 5000, // The start time of the transcription in milliseconds audio_end_at: 15000, // The end time of the transcription in milliseconds }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); }; run(); ``` ```javascript title="JavaScript" highlight={19-20} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web audio_start_from: 5000, // The start time of the transcription in milliseconds audio_end_at: 15000, // The end time of the transcription in milliseconds }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` --- # Transcript Status URL: https://www.assemblyai.com/docs/pre-recorded-audio/check-transcript-status Source: docs/pre-recorded-audio/check-transcript-status.mdx Navigation: Pre-recorded STT > Features > Transcription operations Description: Transcript Status documentation. After you've submitted a file for transcription, your transcript has one of the following statuses: | Status | Description | | ------------ | -------------------------------------------------- | | `processing` | The audio file is being processed. | | `queued` | The audio file is waiting to be processed. This status is only returned when a job is actually queued, such as when you've exceeded your rate limit. Otherwise, transcripts go directly to `processing`. | | `completed` | The transcription has completed successfully. | | `error` | An error occurred while processing the audio file. | ## Handling errors If the transcription fails, the status of the transcript is `error`, and the transcript includes an `error` property explaining what went wrong. ```python title="Python SDK" highlight={11-13} from assemblyai import TranscriptStatus from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcriber = Transcriber(api_key="") transcript = transcriber.transcribe(audio_file, config) if transcript.status == TranscriptStatus.error: print(f"Transcription failed: {transcript.error}") exit(1) print(transcript.text) ``` ```python title="Python" highlight={33-34} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript: {transcription_result['text']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" highlight={16-19} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { console.error(`Transcription failed: ${transcript.error}`); process.exit(1); } console.log(transcript.text); }; run(); ``` ```javascript title="JavaScript" highlight={35-36} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` A transcription may fail for various reasons: - Unsupported file format - Missing audio in file - Unreachable audio URL If a transcription fails due to a server error, we recommend that you resubmit the file for transcription to allow another server to process the audio. A "400 Bad Request" error typically indicates that there's a problem with the formatting or content of the API request. Double-check the syntax of your request and ensure that all required parameters are included as described in the [API reference](/pre-recorded-audio/api-reference/transcripts/submit). If the issue persists, contact our support team for assistance. --- # Transcript export options URL: https://www.assemblyai.com/docs/pre-recorded-audio/export-transcripts-as-srt-vtt-or-text Source: docs/pre-recorded-audio/export-transcripts-as-srt-vtt-or-text.mdx Navigation: Pre-recorded STT > Features > Transcription operations Description: Transcript export options documentation. This page explains the different ways you can export and format your transcript data, including SRT/VTT caption files, paragraphs and sentences, and word-level timestamps. The plain-text transcript returned in the `text` field is raw transcript text only — it does not include timestamps or speaker labels. To build a custom text export in a format such as `[Speaker] timecode -> text`, enable [speaker labels](/pre-recorded-audio/label-speakers) and construct the output from the `utterances` array (which contains `speaker`, `start`, `end`, and `text` for each utterance). See the [Timestamped transcripts](/guides/timestamped-transcripts) guide and the [Create subtitles with speaker labels](/pre-recorded-audio/guides/speaker_labelled_subtitles) cookbook for examples. ## Export SRT or VTT caption files You can export completed transcripts in SRT or VTT format, which can be used for subtitles and closed captions in videos. You can also customize the maximum number of characters per caption by specifying the `chars_per_caption` parameter. ```python title="Python SDK" highlight={13-16} expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcriber = Transcriber(api_key="", config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") srt = transcript.export_subtitles_srt( # Optional: Customize the maximum number of characters per caption chars_per_caption=32 ) with open(f"transcript_{transcript.id}.srt", "w") as srt_file: srt_file.write(srt) # vtt = transcript.export_subtitles_vtt() # with open(f"transcript_{transcript_id}.vtt", "w") as vtt_file: # vtt_file.write(vtt) ``` ```python title="Python" highlight={40} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) # chars_per_caption is optional srt_response = requests.get(f"{polling_endpoint}/srt?chars_per_caption=32", headers=headers) with open(f"transcript_{transcript_id}.srt", "w") as srt_file: srt_file.write(srt_response.text) # vtt_response = requests.get(f"{polling_endpoint}/vtt", headers=headers) # with open(f"transcript_{transcript_id}.vtt", "w") as vtt_file: # vtt_file.write(vtt_response.text) ``` ```javascript title="JavaScript SDK" highlight={18} expandable import { AssemblyAI } from "assemblyai"; import fs from "fs"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); let srt = await client.transcripts.subtitles(transcript.id, "srt", 32); fs.writeFileSync(`transcript_${transcript.id}.srt`, srt); // let vtt = await client.transcripts.subtitles(transcript.id, 'vtt', 32) // fs.writeFileSync(`transcript_${transcript.id}.vtt`, vtt) }; run(); ``` ```javascript title="JavaScript" highlight={43} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } const srtEndpoint = `${baseUrl}/v2/transcript/${transcriptId}/srt?chars_per_caption=32`; res = await fetch(srtEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const srt = await res.text(); fs.writeFileSync(`transcript_${transcriptId}.srt`, srt); // const vttEndpoint = `${baseUrl}/v2/transcript/${transcriptId}/vtt?chars_per_caption=32` // const vtt = await fetch(vttEndpoint, { headers }).then(res => res.text()) // fs.writeFileSync(`transcript_${transcriptId}.vtt`, vtt) ``` ## Export paragraphs You can retrieve transcripts that are automatically segmented into paragraphs. The text of the transcript is broken down by paragraphs, along with additional metadata. ```python title="Python SDK" highlight={13-16} expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcriber = Transcriber(api_key="", config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") paragraphs = transcript.get_paragraphs() for paragraph in paragraphs: print(paragraph.text) print() ``` ```python title="Python" highlight={39-42} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) paragraphs = requests.get(polling_endpoint + '/paragraphs', headers=headers).json()['paragraphs'] for paragraph in paragraphs: print(paragraph['text']) print() ``` ```javascript title="JavaScript SDK" highlight={15-18} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); const { paragraphs } = await client.transcripts.paragraphs(transcript.id); for (const paragraph of paragraphs) { console.log(paragraph.text); } }; run(); ``` ```javascript title="JavaScript" highlight={42-50} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } const paragraphsEndpoint = `${baseUrl}/v2/transcript/${transcriptId}/paragraphs`; res = await fetch(paragraphsEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const paragraphsResponse = await res.json(); const paragraphs = paragraphsResponse.paragraphs; for (const paragraph of paragraphs) { console.log(paragraph.text); console.log(); } ``` ## Export sentences You can retrieve transcripts that are automatically segmented into sentences, for a more reader-friendly experience. The text of the transcript is broken down by sentences, along with additional metadata. ```python title="Python SDK" highlight={13-16} expandable from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcriber = Transcriber(api_key="", config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") sentences = transcript.get_sentences() for sentence in sentences: print(sentence.text) print() ``` ```python title="Python" highlight={39-42} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) sentences = requests.get(polling_endpoint + '/sentences', headers=headers).json()['sentences'] for sentence in sentences: print(sentence['text']) print() ``` ```javascript title="JavaScript SDK" highlight={15-19} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); const { sentences } = await client.transcripts.sentences(transcript.id); for (const sentence of sentences) { console.log(sentence.text); } }; run(); ``` ```javascript title="JavaScript" highlight={44-50} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } const sentencesEndpoint = `${baseUrl}/v2/transcript/${transcriptId}/sentences`; res = await fetch(sentencesEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const sentencesResponse = await res.json(); const sentences = sentencesResponse.sentences; for (const sentence of sentences) { console.log(sentence.text); console.log(); } ``` The response is an array of objects, each representing a sentence or a paragraph in the transcript. See the [API reference](/api-reference/transcripts/get-sentences) for more info. ## Word-level timestamps The response also includes an array with information about each word: ```python title="Python SDK" highlight={10-11} from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcriber = Transcriber(api_key="") transcript = transcriber.transcribe(audio_file, config) for word in transcript.words: print(f"Word: {word.text}, Start: {word.start}, End: {word.end}, Confidence: {word.confidence}") ``` ```python title="Python" highlight={30-31} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': for word in transcription_result['words']: print(f"Word: {word['text']}, Start: {word['start']}, End: {word['end']}, Confidence: {word['confidence']}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" highlight={21} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); // Print word-level details for (const word of transcript.words) { console.log( `Word: ${word.text}, Start: ${word.start}, End: ${word.end}, Confidence: ${word.confidence}` ); } }; run(); ``` ```javascript title="JavaScript" highlight={35-38} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); // Print word-level details for (const word of transcriptionResult.words) { console.log( `Word: ${word.text}, Start: ${word.start}, End: ${word.end}, Confidence: ${word.confidence}` ); } break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` ### API Reference ```js expandable { words: [ { text: "Smoke", start: 240, end: 640, confidence: 0.70473, speaker: null, }, { text: "from", start: 680, end: 968, confidence: 0.99967, speaker: null, }, { text: "hundreds", start: 1024, end: 1416, confidence: 0.99795, speaker: null, }, { text: "of", start: 1448, end: 1592, confidence: 0.99926, speaker: null, }, { text: "wildfires", start: 1616, end: 2248, confidence: 0.99838, speaker: null, }, { text: "in", start: 2264, end: 2440, confidence: 0.99782, speaker: null, }, { text: "Canada", start: 2480, end: 2968, confidence: 0.99977, speaker: null, }, ], } ``` ## Additional resources Learn how to create caption files that include speaker identification. --- # Delete Transcripts URL: https://www.assemblyai.com/docs/pre-recorded-audio/delete-transcripts Source: docs/pre-recorded-audio/delete-transcripts.mdx Navigation: Pre-recorded STT > Features > Transcription operations Description: Delete Transcripts documentation. You can remove the data from the transcript and mark it as deleted. ```python title="Python SDK" highlight={22} expandable import assemblyai from assemblyai.prerecorded.v2 import Transcriber, Transcript, TranscriptionConfig # `Transcript.delete_by_id()` and `Transcript.get_by_id()` use the default client, which reads the global settings. assemblyai.settings.api_key = "" # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( language_detection=True ) transcriber = Transcriber(config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) Transcript.delete_by_id(transcript.id) transcript = Transcript.get_by_id(transcript.id) print(transcript.text) ``` ```python title="Python" highlight={39} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = base_url + "/v2/transcript/" + transcript_id while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) response = requests.delete(polling_endpoint, headers=headers) print(response.json()['text']) ``` ```javascript title="JavaScript SDK" highlight={18} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); const res = await client.transcripts.delete(transcript.id); console.log(res); }; run(); ``` ```javascript title="JavaScript" highlight={42-44} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } res = await fetch(pollingEndpoint, { method: "DELETE", headers, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const deleteResponse = await res.json(); console.log(deleteResponse.text); ``` **Data Retention** See [here](/data-retention-and-model-training) for more information on how long AssemblyAI retains data in the production environment. --- # Upload a media file URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/files/upload Source: docs/pre-recorded-audio/api-reference/files/upload.mdx Navigation: Pre-recorded STT > API reference Description: Upload a media file to AssemblyAI. Metadata: - openapi: ../../../openapi.yaml POST /v2/upload --- # Submit a transcript URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/submit Source: docs/pre-recorded-audio/api-reference/transcripts/submit.mdx Navigation: Pre-recorded STT > API reference Description: Create a transcription request. Metadata: - openapi: ../../../openapi.yaml POST /v2/transcript --- # Get a transcript URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/get Source: docs/pre-recorded-audio/api-reference/transcripts/get.mdx Navigation: Pre-recorded STT > API reference Description: Retrieve a transcript by ID. Metadata: - openapi: ../../../openapi.yaml GET /v2/transcript/{transcript_id} --- # Get transcript sentences URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/get-sentences Source: docs/pre-recorded-audio/api-reference/transcripts/get-sentences.mdx Navigation: Pre-recorded STT > API reference Description: Retrieve transcript sentences. Metadata: - openapi: ../../../openapi.yaml GET /v2/transcript/{transcript_id}/sentences --- # Get transcript paragraphs URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/get-paragraphs Source: docs/pre-recorded-audio/api-reference/transcripts/get-paragraphs.mdx Navigation: Pre-recorded STT > API reference Description: Retrieve transcript paragraphs. Metadata: - openapi: ../../../openapi.yaml GET /v2/transcript/{transcript_id}/paragraphs --- # Get subtitles URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/get-subtitles Source: docs/pre-recorded-audio/api-reference/transcripts/get-subtitles.mdx Navigation: Pre-recorded STT > API reference Description: Export subtitles for a transcript. Metadata: - openapi: ../../../openapi.yaml GET /v2/transcript/{transcript_id}/{subtitle_format} --- # Get redacted audio URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/get-redacted-audio Source: docs/pre-recorded-audio/api-reference/transcripts/get-redacted-audio.mdx Navigation: Pre-recorded STT > API reference Description: Retrieve redacted audio for a transcript. Metadata: - openapi: ../../../openapi.yaml GET /v2/transcript/{transcript_id}/redacted-audio --- # Word search URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/word-search Source: docs/pre-recorded-audio/api-reference/transcripts/word-search.mdx Navigation: Pre-recorded STT > API reference Description: Search a transcript for words or phrases. Metadata: - openapi: ../../../openapi.yaml GET /v2/transcript/{transcript_id}/word-search --- # List transcripts URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/list Source: docs/pre-recorded-audio/api-reference/transcripts/list.mdx Navigation: Pre-recorded STT > API reference Description: List transcripts. Metadata: - openapi: ../../../openapi.yaml GET /v2/transcript --- # Delete a transcript URL: https://www.assemblyai.com/docs/pre-recorded-audio/api-reference/transcripts/delete Source: docs/pre-recorded-audio/api-reference/transcripts/delete.mdx Navigation: Pre-recorded STT > API reference Description: Delete a transcript. Metadata: - openapi: ../../../openapi.yaml DELETE /v2/transcript/{transcript_id} --- # Webhooks for pre-recorded audio URL: https://www.assemblyai.com/docs/pre-recorded-audio/webhooks Source: docs/pre-recorded-audio/webhooks.mdx Navigation: Pre-recorded STT > Advanced Description: Get notified when a pre-recorded audio transcription is ready. Webhooks are custom HTTP callbacks that you can define to get notified when your pre-recorded audio transcripts are ready. This guide covers webhooks for [pre-recorded audio transcription](/pre-recorded-audio). For webhooks with streaming audio, see [Webhooks for streaming speech-to-text](/streaming/webhooks). To use webhooks, you need to set up your own webhook receiver to handle webhook deliveries. ## Create a webhook for a transcription **Don't have a webhook endpoint yet?** Create a test webhook endpoint with [webhook.site](https://webhook.site) to test your webhook integration. To create a webhook, set the `webhook_url` parameter when you create a new transcription. The URL must be accessible from AssemblyAI's servers. ```python expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, "webhook_url": "https://example.com/webhook" } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) ``` To create a webhook, use `set_webhook()` on the transcription config. The URL must be accessible from AssemblyAI's servers. Use `submit()` instead of `transcribe()` to create a transcription without waiting for it to complete. ```python from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig().set_webhook("https://example.com/webhook") transcriber = Transcriber(api_key="") transcriber.submit(audio_file, config) ``` To create a webhook, set the `webhook_url` parameter when you create a new transcription. The URL must be accessible from AssemblyAI's servers. ```javascript expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, webhook_url: "https://example.com/webhook", }; const url = `${baseUrl}/v2/transcript`; const response = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!response.ok) throw new Error(`Error: ${response.status}`); ``` To create a webhook, include the `webhook_url` parameter when you create a new transcription. The URL must be accessible from AssemblyAI's servers. Use `submit()` instead of `transcribe()` to create a transcription without waiting for it to complete. ```javascript import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, webhook_url: "https://example.com/webhook", }; const run = async () => { const transcript = await client.transcripts.submit(params); }; run(); ``` ## Handle webhook deliveries When the transcript is ready, AssemblyAI will send a `POST` HTTP request to the URL that you specified. Your webhook endpoint must return a 2xx HTTP status code within 10 seconds to indicate successful receipt. If a 2xx status is not received within 10 seconds, AssemblyAI will retry the webhook call up to a total of 10 attempts. If at any point your endpoint returns a 4xx status code, the webhook call is considered failed and will not be retried. **Static Webhook IP addresses** AssemblyAI sends all webhook deliveries from fixed IP addresses: | Region | IP Address | | ------ | -------------- | | US | `44.238.19.20` | | EU | `54.220.25.36` | ### Delivery payload The webhook delivery payload contains a JSON object with the following properties: ```json { "transcript_id": "5552493-16d8-42d8-8feb-c2a16b56f6e8", "status": "completed" } ``` | Key | Type | Description | | --------------- | ------ | ------------------------------------------------------------ | | `transcript_id` | string | The ID of the transcript. | | `status` | string | The status of the transcript. Either `completed` or `error`. | The webhook payload only contains the `transcript_id` and `status`. It doesn't include the transcript text or error details. If the `status` is `"error"`, make a `GET /v2/transcript/{transcript_id}` request to retrieve the error message from the `error` field in the response. See [Retrieve a transcript with the transcript ID](#retrieve-a-transcript-with-the-transcript-id) below. ### Webhook status code When you retrieve the transcript using the transcript ID, the JSON response will include a `webhook_status_code` field. This field contains the HTTP status code that AssemblyAI received when calling your webhook endpoint, allowing you to verify that the webhook was delivered successfully. ### Retrieve a transcript with the transcript ID ```python import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } transcript_id = "" polling_endpoint = f"https://api.assemblyai.com/v2/transcript/{transcript_id}" transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") ``` ```python import assemblyai from assemblyai.prerecorded.v2 import Transcript # `Transcript.get_by_id()` uses the default client, which reads the global settings. assemblyai.settings.api_key = "" transcript = Transcript.get_by_id("") if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```javascript import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const transcriptId = ""; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; let res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } ``` ```javascript import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); const transcript = await client.transcripts.get(""); if (transcript.status === "error") { throw new Error(`Transcription failed: ${transcript.error}`); } console.log(transcript.text); ``` ## Authenticate webhook deliveries You can authenticate webhook deliveries from AssemblyAI by including a custom HTTP header in the request. To add an authentication header, include the auth header name and value in `set_webhook()`. ```python {2} config = TranscriptionConfig().set_webhook( "https://example.com/webhook", "X-My-Webhook-Secret", "secret-value" ) transcriber = Transcriber(api_key="") transcriber.submit(audio_url, config) ``` To add an authentication header, include the `webhook_auth_header_name` and `webhook_auth_header_value` parameters. ```javascript {5-6} client.transcripts.submit({ audio: 'https://assembly.ai/wildfires.mp3', webhook_url: 'https://example.com/webhook', webhook_auth_header_name: "X-My-Webhook-Secret", webhook_auth_header_value: "secret-value" }) ``` To add an authentication header, include the auth header name and value in `set_webhook()`. ```python expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, "webhook_url": "https://example.com/webhook", "webhook_auth_header_name": "X-My-Webhook-Secret", "webhook_auth_header_value": "secret-value" } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) ``` To add an authentication header, include the auth header name and value in `set_webhook()`. ```python from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig().set_webhook("https://example.com/webhook", "X-My-Webhook-Secret", "secret-value") transcriber = Transcriber(api_key="") transcriber.submit(audio_file, config) ``` To add an authentication header, include the `webhook_auth_header_name` and `webhook_auth_header_value` parameters. ```javascript expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, webhook_url: "https://example.com/webhook", webhook_auth_header_name: "X-My-Webhook-Secret", webhook_auth_header_value: "secret-value", }; const url = `${baseUrl}/v2/transcript`; const response = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!response.ok) throw new Error(`Error: ${response.status}`); ``` To add an authentication header, include the `webhook_auth_header_name` and `webhook_auth_header_value` parameters. ```javascript expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, webhook_url: "https://example.com/webhook", webhook_auth_header_name: "X-My-Webhook-Secret", webhook_auth_header_value: "secret-value", }; const run = async () => { const transcript = await client.transcripts.submit(params); }; run(); ``` ## Add metadata to webhook deliveries To associate metadata for a specific transcription request, you can add your own query parameters to the webhook URL. ```plain https://example.com/webhook?customer_id=1234&order_id=5678 ``` Now, when you receive the webhook delivery, you'll know the customer who requested it. ## Failed webhook deliveries Webhook deliveries can fail for multiple reasons. For example, if your server is down or takes more than 10 seconds to respond. If a webhook delivery fails, AssemblyAI will attempt to redeliver it up to 10 times, waiting 10 seconds between each attempt. If all attempts fail, AssemblyAI considers the delivery as permanently failed. --- # Benchmarks URL: https://www.assemblyai.com/docs/pre-recorded-audio/benchmarks Source: docs/pre-recorded-audio/benchmarks.mdx Navigation: Pre-recorded STT > Advanced Description: Industry-leading accuracy for pre-recorded speech-to-text. Benchmarks are an important first step before running your own [evaluation](/pre-recorded-audio/evaluations). Below are the current benchmarks for our pre-recorded models so you can assess performance across accuracy, latency, and error rates. Public benchmarks can be misleading due to overfitting and benchmark gaming. We strongly recommend [running your own evaluation](/pre-recorded-audio/evaluations) on your audio data to identify the best model for your use case. For the full interactive benchmark experience with competitive comparisons, visit [assemblyai.com/benchmarks](https://www.assemblyai.com/benchmarks). ## Word error rate (WER) Word Error Rate (WER) is the classical metric for speech-to-text accuracy. It counts substitutions, deletions, and insertions against a reference transcript, divided by the total word count in the ground truth. AssemblyAI Universal-3.5 Pro achieves a mean WER of **5.6%** (median **4.9%**) on English benchmarks. WER weights every word equally, so a misrecognized filler word counts the same as a misrecognized email address or medication name. For production voice workflows, we recommend pairing WER with [Missed entity rate](#missed-entity-rate), which measures accuracy on the high-stakes entities — names, emails, phone numbers, and medical terms — that actually drive end-user outcomes. ### English benchmarks Most recent update: **January 2026**. | Dataset | Universal-3.5 Pro WER (%) | Universal-2 WER (%) | Relative gain vs Universal-2 | | ------------------------------------------------------------------------------------------ | ---------------------------------- | ---------------------------------- | ------------------------------------- | | **Overall Performance** | **Mean: 5.6%** \| **Median: 4.9%** | **Mean: 6.1%** \| **Median: 6.5%** | **Mean: 8.2%** \| **Median: 24.6%** | | [commonvoice](https://github.com/common-voice/cv-dataset) | 4.87% | 6.48% | 24.8% | | [earnings21](https://huggingface.co/datasets/distil-whisper/earnings21) | 8.80% | 9.37% | 6.1% | | [librispeech_test_clean](https://huggingface.co/datasets/AudioLLMs/librispeech_test_clean) | 1.52% | 1.68% | 9.5% | | [librispeech_test_other](https://huggingface.co/datasets/openslr/librispeech_asr) | 2.69% | 3.00% | 10.3% | | [meanwhile](https://huggingface.co/datasets/distil-whisper/meanwhile) | 4.22% | 4.41% | 4.3% | | [tedlium](https://huggingface.co/datasets/lmms-lab/tedlium) | 6.77% | 7.30% | 7.3% | | [rev16](https://huggingface.co/datasets/distil-whisper/rev16) | 10.29% | 10.32% | 0.3% | ### Multilingual benchmarks Most recent update: **January 2026**. Dataset: [FLEURS](https://huggingface.co/datasets/google/fleurs). | Language Code | Language | Universal-3.5 Pro WER (%) | Universal-2 WER (%) | Relative gain vs Universal-2 | | ------------- | ---------- | ----------------------- | ------------------- | ---------------------------- | | **Average** | **All** | **4.58%** | **7.42%** | **38.3%** | | de | German | 4.88% | 6.22% | 21.5% | | en | English | - | 4.38% | - | | es | Spanish | 3.98% | 4.56% | 12.7% | | fi | Finnish | - | 10.10% | - | | fr | French | 4.98% | 7.56% | 34.1% | | hi | Hindi | - | 7.38% | - | | it | Italian | 3.69% | 4.75% | 22.3% | | ja | Japanese | - | 7.79% | - | | ko | Korean | - | 14.54% | - | | nl | Dutch | - | 7.79% | - | | pl | Polish | - | 6.63% | - | | pt | Portuguese | 5.39% | 5.98% | 9.9% | | ru | Russian | - | 5.80% | - | | tr | Turkish | - | 8.12% | - | | uk | Ukrainian | - | 7.42% | - | | vi | Vietnamese | - | 9.75% | - | ## Missed entity rate For production voice workflows, the actual words that matter most are entities — names, organizations, emails, phone numbers, and medical terms. The Missed Entity Rate (MER) measures how often a model fails to correctly transcribe these high-stakes terms. See [Missed Entity Rate](/pre-recorded-audio/evaluations#missed-entity-rate-mer) for the full definition. Universal-3.5 Pro delivers relative gains over Universal-2 across every entity category we track for voice workflows, with the largest improvements on emails, locations, and medical terms. | Entity type | Universal-3.5 Pro MER (%) | Universal-2 MER (%) | Relative gain vs Universal-2 | | ------------------- | ----------------------- | ------------------- | ---------------------------- | | Medical terms | 13.15% | 18.43% | 28.6% | | Locations | 8.61% | 12.40% | 30.6% | | Job titles | 9.03% | 9.86% | 8.4% | | Organization names | 17.02% | 20.96% | 18.8% | | Email addresses | 33.76% | 53.81% | 37.3% | | Phone numbers | 13.14% | 14.69% | 10.6% | | Credit card numbers | 21.83% | 25.07% | 12.9% | ## Hallucinations and consecutive errors Hallucinations are a critical concern in production STT systems. AssemblyAI reduces hallucinations by 30% compared to Whisper, across three error categories: - **Fabrications** — words inserted that were never spoken - **Omissions** — spoken words that are missing from the transcript - **Hallucinations** — extended sequences of fabricated content ## Benchmark challenges Models are often trained on publicly available datasets — sometimes the very same datasets used for evaluation. When this happens, the model becomes overfit to the evaluation set and will show artificially strong performance on standard WER tests. This makes WER potentially misleading, as real-world performance on unseen audio will be significantly worse. ## External benchmarks For third-party benchmarks, we recommend the [Hugging Face ASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard). Note that many models listed require self-hosting and lack production features like speaker diarization and automatic language detection. ## Methodology Our benchmarks are evaluated across **250+ hours of audio data**, **80,000+ audio files**, and **26 datasets**. We apply standard text normalization before calculating metrics. For full details on our methodology, visit [assemblyai.com/benchmarks](https://www.assemblyai.com/benchmarks). ## Run your own benchmark We'd be happy to help. AssemblyAI has a benchmarking tool to help you run a custom evaluation against your real audio files. [Contact us](https://www.assemblyai.com/contact/support) for more information. You can also run your own benchmarks following the [Hugging Face framework](https://github.com/huggingface/open_asr_leaderboard) which provides a GitHub repo with full instructions. --- # Evaluating Pre-recorded STT models URL: https://www.assemblyai.com/docs/pre-recorded-audio/evaluations Source: docs/pre-recorded-audio/evaluations.mdx Navigation: Pre-recorded STT > Advanced Description: Evaluating Pre-recorded STT models documentation. ## Introduction The high level objective of a pre-recorded STT model evaluation is to answer the question: _Which Speech-to-text model is the best for my product?_ This guide provides a step-by-step framework for evaluating and benchmarking pre-recorded Speech-to-text models, with specific guidance for evaluating Universal-3.5 Pro and its prompting capabilities. Need help with evaluations or prompt optimization? [Contact our Sales team](https://www.assemblyai.com/contact/sales) — we can help you design an evaluation, optimize prompts for your audio, and benchmark against your ground truth data. ## Evaluation metrics ### Traditional metrics #### Word Error Rate (WER) $$ WER = \frac{S + D + I}{N} $$ This formula takes the number of Substitutions (S), Deletions (D), and Insertions (I), and divides their sum by the Total Number of Words in the ground truth transcript (N). While WER calculation may seem simple, it requires a methodical granular approach and reliable reference data. Word Error Rate can tell you how “different” the automatic transcription was compared to the human transcription, and *generally*, this is a reliable metric to determine how “good” a transcription is. For more info on WER as a metric, read Dylan Fox’s [blog post on Word Error Rate](https://www.assemblyai.com/blog/word-error-rate). #### Concatenated minimum-Permutation Word Error Rate (cpWER) $$ \text{cpWER} = \frac{S_{\text{spk}} + D + I}{N} $$ cpWER is similar to WER, but it also measures the number of errors a speech recognition model makes where words with incorrectly-ascribed speakers are considered to be incorrect. The primary difference from standard WER is how _S_ is calculated: $S_{\text{spk}}$ counts both word substitutions and correctly transcribed words that are assigned to the wrong speaker. A correct word with an incorrect speaker label counts as a substitution error, thereby penalizing both transcription and speaker diarization mistakes. #### Formatted WER (F-WER) F-WER is similar to WER but F-WER does not apply text normalization, so all formatting differences are accounted for, in addition to word differences when computing the WER. Therefore, F-WER is always higher than or equal to WER. #### Sentence Error Rate (SER) $$ \text{SER} = \frac{N_{\text{err}}}{N_{\text{sent}}} $$ The Sentence Error Rate (SER) is the ratio of the number of sentences with one or more errors to the total number of sentences. #### Diarization Error Rate (DER) $$ DER = \frac{false alarm + missed detection + confusion}{total} $$ This formula takes the duration of non-speech incorrectly classified as speech (false alarm), the duration of speech incorrectly classified as non-speech (missed detection), the duration of speaker confusion (confusion), and divides the sum over the total speech duration. #### Missed Entity Rate (MER) $$ \text{MER} = 1 - \frac{N_{\text{rec}}}{N_{\text{total}}} $$ Fundamentally, MER is a negative recall rate computed for specified target entities. It is defined as the number of correctly transcribed entities relative to their total occurrence count. It accounts for multiple occurrences of the same entity and their positions within the hypothesis transcription. Our Research team proposes this as the best metric to measure the effectiveness of word boost. For a simpler approach when evaluating high-stakes entities (such as credit card numbers, names, or dosages), consider a binary pass/fail metric per file: score each file as 1 if the model captured the target entity correctly, or 0 if it did not. Run this across 100 or more files for statistical reliability. This is especially useful when you need a quick signal on entity accuracy at scale without computing full MER breakdowns. ### Metrics for Universal-3.5 Pro Universal-3.5 Pro is significantly more capable than prior models, and traditional WER alone may not fully capture its performance. The following metrics provide a more complete picture. #### Semantic WER Traditional WER treats every difference between the model output and a reference transcript as an error—even when the difference is semantically equivalent. Semantic WER corrects this by normalizing equivalent words and phrases before calculating WER, so that differences like `dr.` vs `doctor` or `1300` vs `thirteen hundred` aren't counted as errors. ##### Rule-based normalization At its simplest, Semantic WER is a preprocessing step. Before running standard WER, apply find-and-replace rules to both the reference and hypothesis transcripts: - **Number formats**: `1300` → `thirteen hundred`, `$5` → `five dollars` - **Abbreviations and titles**: `dr.` → `doctor`, `mr.` → `mister`, `govt` → `government` - **Contractions**: `gonna` → `going to`, `can't` → `cannot` - **Variant spellings**: `grey` → `gray`, `cancelled` → `canceled` - **Filler words**: Remove `um`, `uh`, `you know` from both sides (or keep both—just be consistent) This alone eliminates a significant portion of false errors and can be implemented in a few lines of Python. No model inference required. ##### LLM-based scoring For cases where simple rules can't capture the nuance—was an omission meaningful? Is a proper noun misspelling close enough?—an LLM can perform word-level alignment and classify each difference by severity: - **No penalty**: Semantically equivalent forms (number formats, contractions, variant spellings) - **Minor penalty**: Single-character misspellings, minor grammatical markers - **Major penalty**: Incorrect substitutions, meaning-altering errors, significant omissions or additions of content words These approaches are particularly valuable for Universal-3.5 Pro because the model often transcribes audio more accurately than human transcribers, producing differences that are correct but would be penalized by traditional WER. For an implementation of Semantic WER using Bayesian optimization, see [prompt-seeker](https://github.com/AssemblyAI-Solutions/prompt-seeker). #### LASER score (LLM-based ASR Evaluation Rubric) LASER is a published LLM-based evaluation metric ([Parulekar & Jyothi, EMNLP 2025](https://aclanthology.org/2025.emnlp-main.1257/)) that uses an LLM prompt with detailed examples to classify ASR errors and compute a score: $$ \text{LASER} = 1 - \frac{\text{Total Penalty}}{\text{Reference Word Count}} $$ The LLM aligns each word in the ASR output against the reference transcription and assigns a penalty per word pair: - **No penalty (0)**: Acceptable variations including numerical format differences, abbreviations, compound word splits, transliterations, alternate spellings, proper noun variants, and colloquial terms - **Minor penalty (0.5)**: Small spelling errors (single character) or minor grammatical errors (gender, tense, number markers) that preserve sentence meaning - **Major penalty (1.0)**: Incorrect word substitutions, significant omissions or additions, and reordering that changes meaning LASER provides structured per-error feedback alongside the score. This makes it useful for prompt optimization workflows where you need to understand _why_ a prompt performed poorly, not just _how much_ error there was. For an implementation of LASER scoring, see [aai-cli](https://github.com/alexkroman/aai-cli). ### Why new metrics matter Traditional WER treats every difference between the model output and human transcription as an error. Universal-3.5 Pro’s contextual awareness means it will often transcribe words that human transcribers missed entirely. In traditional WER, these show up as **insertions** (penalized errors), even though the model is correct. This makes WER an unreliable metric when used alone — your evaluation is only as good as your ground truth labels. WER is only as good as your ground truth labels. Human transcriptions contain systematic errors — missed filler words, incorrect proper nouns, simplified speech patterns, and translated code-switching. When Universal-3.5 Pro transcribes audio more accurately than the human label, those improvements show up as WER _errors_. Before reporting WER, manually audit at least 20 insertions to determine what percentage are true errors versus ground truth omissions. In our testing, the majority of insertions were cases where Universal-3.5 Pro correctly transcribed audio that the human transcriber missed. This is why [Artificial Analysis](https://artificialanalysis.ai/speech-to-text), an independent AI benchmarking organization, had to create proprietary evaluation datasets with manually corrected ground truths when building their [Speech-to-Text leaderboard](https://artificialanalysis.ai/speech-to-text). Existing public datasets contain systematic human transcription errors that penalize models which are actually more accurate. ## The evaluation process This section provides a step-by-step guide on how to run an evaluation. The evaluation process should closely match your production environment, including the files you intend to transcribe, the model you intend to use, and the settings applied to those models. ### Step 1: Prepare your evaluation dataset Ensure that the files you use to benchmark are representative of the files you plan to use in production. For example, if you plan to transcribe meetings, gather a set of meeting recordings. If you plan to transcribe phone calls, focus on finding phone calls that match your customer base’s language and region. We recommend using at least 25 files that are representative of your use case. Length is less important than diversity of audio conditions — a good evaluation set covers the range of speakers, accents, audio quality, and vocabulary your model will encounter in production. Then, gather human-labeled data to act as your source of ground truth. Ground truth is accurately transcribed audio data that will serve as the "correct answer" for our benchmark. Human-labeled data can be purchased from an external vendor or created manually. Open-source audio corpora (for example, datasets on [Hugging Face](https://huggingface.co/datasets?task_categories=automatic-speech-recognition)) can serve as a starting point for building ground truth, but they require review and correction before use in production evaluations. These datasets contain the same systematic human transcription errors described below — missing filler words, incorrect proper nouns, and simplified speech patterns — and should be audited against your actual audio before benchmarking. #### Ground truth quality The quality of your ground truth data directly affects the reliability of your evaluation. With Universal-3.5 Pro, this is more important than ever because the model frequently outperforms human transcribers. Common issues with ground truth data: - **Missing filler words**: Human transcribers often omit `um`, `uh`, `like`, and other disfluencies - **Incorrect proper nouns**: Rare names, technical terms, and domain vocabulary are often misspelled - **Simplified speech patterns**: Human transcribers tend to “clean up” speech, missing repetitions, false starts, and self-corrections - **Code-switching errors**: Multilingual segments are frequently translated to English rather than transcribed as spoken Before running evaluations, audit a sample of your ground truth files by listening to the audio and comparing. If your ground truth contains systematic errors, your WER numbers will be misleading. To inspect and correct issues in your ground truth files, use the [Truth File Corrector](https://www.assemblyai.com/dashboard/home) in the AssemblyAI Dashboard (found at the bottom of the left sidebar), which lets you listen back to audio and fix human transcription errors by clicking through differences. See [this article](https://www.assemblyai.com/blog/new-word-error-rate-wer-benchmark) to learn more about why your word error rate (WER) benchmark might be lying to you. #### Dataset diversity A prompt that performs well overall may underperform on specific audio types. Include a diverse mix of audio in your evaluation set and track per-dataset breakdowns: | Audio type | Characteristics | Typical WER range | | --------------------- | --------------------------------------- | -------------------------------------- | | Earnings calls | Clean English, formal vocabulary | Low | | Meeting recordings | Multi-speaker, informal | Moderate | | Code-switching audio | Mixed languages (e.g., English/Spanish) | Higher (normalization affects scoring) | | Medical consultations | Clinical vocabulary, accented speech | Moderate | | Phone calls | Compression artifacts, background noise | Moderate to high | ### Step 2: Establish a baseline Before optimizing prompts, measure your baseline performance by transcribing your evaluation set with Universal-3.5 Pro and **no custom prompt**. The built-in default is already applied when `prompt` is omitted, and it outperforms most custom prompts. Record both WER and Semantic WER so you can track improvements as you layer instructions on top. ### Step 3: Transcribe and evaluate with prompts Transcribe your files using [AssemblyAI’s API](/pre-recorded-audio/getting-started/transcribe-an-audio-file) with Universal-3.5 Pro and your candidate prompts. When crafting evaluation prompts, use the [prompting guide](/pre-recorded-audio/universal-3-5-pro/prompting) as a reference. Key principles: - **Use authoritative language**: The model responds better to `Mandatory:`, `Required:`, and `Always:` than soft language like `try to` or `please` - **Be specific about speech patterns**: Enumerate what you want preserved (disfluencies, filler words, hesitations, repetitions, stutters, false starts, colloquialisms) - **Give instructions, not just context**: `This is a doctor-patient visit. Prioritize accurately transcribing medications and diseases wherever possible.` is far more effective than `This is a doctor-patient visit.` - **Start with fewer instructions, add one at a time**: Every added instruction risks conflicting with another. Add a single instruction, evaluate it against your dataset, and only then add the next. ### Step 4: Text normalization Before calculating WER metrics, both reference (ground truth) and hypothesis (model generated) texts need to be normalized to ensure a fair comparison. This accounts for differences in: - Punctuation and capitalization - Number formatting (e.g., “twenty-one” vs. “21”) - Contractions and abbreviations - Other stylistic variations that don’t affect meaning Normalization can be done with a library like [Whisper Normalizer](https://pypi.org/project/whisper-normalizer/). If you are prompting Universal-3.5 Pro to include `[unclear]` or `[masked]` tags for uncertain audio, ensure your normalizer strips these tags before computing WER. Otherwise, they will be counted as insertions. ### Step 5: Compare and calculate Calculate the error rates using the formulas above or consider using a library like [jiwer](https://github.com/jitsi/jiwer). For Semantic WER, apply text normalization replacements before calculating WER. For LASER scoring, use an LLM-based evaluator (see [Open-source tools](#open-source-tools) below). When reviewing results: - Check per-dataset breakdowns, not just aggregate WER - Audit insertions manually by listening to the audio - Compare both traditional WER and Semantic WER to get a full picture - Track which prompt components improve which audio types ### Qualitative analysis Quantitative metrics don’t capture everything. Qualitative analysis helps you identify differences between STT providers that metrics might miss — for example, how certain key terms are transcribed can make or break a transcript, even if the rest of the transcript has a lower overall error rate. Qualitative analysis is also useful for tie-breaking when benchmarking metrics don’t clearly favor one model over another. Since you’re comparing models against each other, ground truth files aren’t required. **Side-by-side comparison**: Have users compare and pick their preferred transcript between two **formatted** outputs from different STT providers. Tools like [Diffchecker](https://www.diffchecker.com/) or any side-by-side interface work well for this. **LLM as judge**: An LLM can automatically identify differences between two transcriptions and pick a winner. However, be cautious: an LLM judge can be misled by outputs that _look_ correct but contain subtle errors (such as translated code-switching segments that read well in English but don’t reflect what was actually spoken). Always pair LLM-based judgments with spot-checking against the actual audio. **A/B testing in production**: Serve transcripts from different providers to users and collect feedback. You can ask users to score transcripts directly, or track indirect signals like the number of support ticket complaints about transcription quality. ### Domain-specific evaluation considerations WER is not always the right primary metric. Some domains prioritize output qualities that traditional accuracy metrics do not capture: - **Medical scribes**: Customers often evaluate based on **user preference rate** and **readability** — whether clinicians prefer the transcript output for generating clinical notes. Formatting quality, medical terminology accuracy, and structured output can matter more than raw WER. See the [Medical Scribe guides](/medical-scribe-best-practices) for domain-specific evaluation guidance. - **Legal transcription**: Verbatim accuracy including disfluencies and speaker attribution may be more important than clean, readable output. - **Media and entertainment**: Proper noun accuracy for names, places, and brands can outweigh overall WER. When running evaluations for domain-specific use cases, define your success criteria before choosing metrics. If your end users care about readability and preference, include qualitative evaluation (side-by-side comparisons, user preference scoring) alongside quantitative metrics. ## Iterating on prompts Finding the optimal prompt for your use case is an iterative process. There are two main approaches: ### Manual iteration 1. Start with the [default system prompt](/pre-recorded-audio/universal-3-5-pro/prompting#best-all-around-default) or one of the reference prompts below 2. Transcribe a representative sample of your audio 3. Review the output against your ground truth, focusing on the types of errors that matter most for your use case 4. Adjust the prompt to address specific error patterns 5. Re-evaluate and compare ### Automated optimization For large-scale prompt optimization, consider using one of the open-source tools described below. These tools systematically test prompt component combinations and score them against your evaluation data, converging on the best prompt for your specific audio. ### Reference prompts for evaluation Use these prompts directly from the [prompting guide](/pre-recorded-audio/universal-3-5-pro/prompting#recommended-prompts) as your evaluation prompts. #### Evaluation prompt Start with the built-in default ([Best all around](/pre-recorded-audio/universal-3-5-pro/prompting#best-all-around-default)). Omit the `prompt` parameter to use it — you don’t need to set it explicitly: ```text Transcribe with context and proper nouns preserved, where speech is present in the audio. Each language as spoken. English as English. Non-native speakers. ``` For maximum verbatim capture and multilingual code-switching, use the [Verbatim with multilingual support](/pre-recorded-audio/universal-3-5-pro/prompting#verbatim-with-multilingual-support) prompt instead. The trade-off is that the model may occasionally hallucinate disfluencies or language switches that don’t exist in the audio. #### Comparison prompt (for identifying model uncertainty) This is the [Handling unclear audio with `[unclear]`](/pre-recorded-audio/universal-3-5-pro/prompting#handling-unclear-audio-with-unclear) prompt. Run it alongside the evaluation prompt on the same audio and diff the outputs to find where the model is guessing: ```text Always: Transcribe speech exactly as heard. If uncertain or audio is unclear, mark as [unclear]. After the first output, review the transcript again. Pay close attention to hallucinations, misspellings, or errors, and revise them like a computer performing spell and grammar checks. Ensure words and phrases make grammatical sense in sentences. ``` By comparing the two outputs, you can identify exactly which segments the model is least confident about. This is useful for: - Evaluating how the model handles unclear or noisy audio - Finding segments where the model’s guesses may be incorrect - Prioritizing which audio segments to manually review - Understanding whether WER differences are coming from genuine errors or uncertain segments ### What works and what doesn’t The authoritative list lives in the prompting guide — see [What works / what to avoid](/pre-recorded-audio/universal-3-5-pro/prompting#what-works--what-to-avoid). The same rules apply when building evaluation prompts: lead with `Transcribe…`, use authoritative language (`Required:`, `Mandatory:`, `Always:`), describe the _pattern_ to watch for, and add instructions one at a time. Listing specific word examples in your prompt causes hallucinations. The model becomes over-eager to insert those exact words into the transcript, even when they weren’t spoken. For example, `Pharmaceutical accuracy required (omeprazole over omeprizole, metformin over metforman)` will cause the model to hallucinate those drug names. Instead, describe the _pattern_ of entities to prioritize: `Pharmaceutical accuracy required across all medications and drug names`. If you already know the specific terms, use [keyterms prompting](/pre-recorded-audio/universal-3-5-pro/prompting#keyterms-prompting) instead — it’s optimized for term boosting and more reliable than describing terms in a free-form prompt. See the [prompting guide](/pre-recorded-audio/universal-3-5-pro/prompting#entity-accuracy-and-spelling) for more details. ## Open-source tools ### aai-cli [aai-cli](https://github.com/alexkroman/aai-cli) is a command-line tool for evaluating and optimizing transcription prompts. It supports: - **Prompt evaluation**: Score a prompt against datasets from Hugging Face or your own audio files using WER and LASER metrics - **Prompt optimization**: Automatically iterate on prompts using DSPy GEPA with LASER feedback, where an LLM reflects on transcription errors and proposes improved prompts - **Dataset discovery**: Search and load audio datasets from Hugging Face for benchmarking ```bash # Evaluate a prompt aai eval --prompt "Transcribe verbatim." --max-samples 50 # Optimize a prompt aai optimize --starting-prompt "Transcribe verbatim." --iterations 5 --samples 50 ``` ### prompt-seeker [prompt-seeker](https://github.com/AssemblyAI-Solutions/prompt-seeker) uses Bayesian optimization (Optuna TPE) to systematically find the best transcription prompt by testing component combinations across diverse audio datasets and scoring with Semantic WER. It supports: - **Component-based optimization**: Modular prompt pieces (language, disfluency, punctuation, etc.) are tested in combinations - **Meta-optimization**: An LLM designs new component spaces between optimization rounds based on accumulated findings - **Per-dataset analysis**: Breakdown of what works for each audio type in your evaluation set ```bash # Run optimization (50 trials across your data) uv run python -m prompt_seeker.cli optimize \ --datasets "my_calls:50" \ --trials 50 -c 20 # Run the meta-optimizer (Claude designs rounds autonomously) uv run python -m prompt_seeker.cli meta-optimize \ --datasets "my_calls:50" \ --rounds 3 --trials 50 -c 10 ``` Both tools require ground truth transcriptions for scoring. If you don’t have ground truth yet, transcribe a sample of your audio manually and use that as your starting point. ## Conclusion Evaluating Universal-3.5 Pro requires going beyond traditional WER. The model’s contextual awareness and prompting capabilities mean that evaluation is as much about finding the right prompt as it is about measuring accuracy. Use Semantic WER or LASER alongside traditional WER, audit your ground truth data carefully, and iterate on prompts systematically to find the best configuration for your audio. --- # Cloud Endpoints and Data Residency URL: https://www.assemblyai.com/docs/pre-recorded-audio/select-the-region Source: docs/pre-recorded-audio/select-the-region.mdx Navigation: Pre-recorded STT > Advanced Description: Cloud Endpoints and Data Residency documentation. Choose the endpoint that best fits your application's requirements—whether that's using the default US region or ensuring your audio data stays within the European Union. ## Default Endpoint The default endpoint (`api.assemblyai.com`) processes your pre-recorded audio transcription requests in the US region. If you don't specify a base URL, this is the endpoint used for the request. ## EU Data Residency The EU endpoint (`api.eu.assemblyai.com`) guarantees your data never leaves the European Union. This is designed for organizations with strict data residency and governance requirements—your audio and transcription data will remain entirely within the EU. ## Endpoints | Endpoint | Base URL | Description | | ------------ | ----------------------- | -------------------- | | US (default) | `api.assemblyai.com` | Data stays in the US | | EU | `api.eu.assemblyai.com` | Data stays in the EU | ## Which endpoint should I use? - **No data residency requirements?** Use the **default endpoint**. No configuration change is needed. - **Need EU data residency?** Use the **EU endpoint** to ensure your audio and transcription data stays within the European Union. ## How to use it Update your base URL to your preferred endpoint. Select an endpoint tab below to see examples for each. The US endpoint is the default. If you're using the SDKs without overriding the base URL, you're already using this endpoint. No configuration change is required. ```python title="Python SDK" from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig # No base_url override needed — US is the default # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, ) transcriber = Transcriber(api_key="", config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" highlight={4} expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "speech_models": ["universal-3-5-pro", "universal-2"], "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = f"{base_url}/v2/transcript/{transcript_id}" while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" expandable import { AssemblyAI } from "assemblyai"; // No baseUrl override needed — US is the default const client = new AssemblyAI({ apiKey: "", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, speech_models: ["universal-3-5-pro", "universal-2"], language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); }; run(); ``` ```javascript title="JavaScript" highlight={4} expandable import fs from "fs-extra"; const baseUrl = "https://api.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web speech_models: ["universal-3-5-pro", "universal-2"], language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` Use the EU endpoint to keep your audio and transcription data within the European Union. ```python title="Python SDK" highlight={7} from assemblyai import Client, Settings from assemblyai.prerecorded.v2 import Transcriber, TranscriptionConfig client = Client( settings=Settings( api_key="", base_url="https://api.eu.assemblyai.com", ) ) # audio_file = "./local_file.mp3" audio_file = "https://assembly.ai/wildfires.mp3" config = TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, ) transcriber = Transcriber(client=client, config=config) transcript = transcriber.transcribe(audio_file) if transcript.status == "error": raise RuntimeError(f"Transcription failed: {transcript.error}") print(transcript.text) ``` ```python title="Python" highlight={4} expandable import requests import time base_url = "https://api.eu.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "speech_models": ["universal-3-5-pro", "universal-2"], "language_detection": True } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = f"https://api.eu.assemblyai.com/v2/transcript/{transcript_id}" while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': print(f"Transcript ID: {transcript_id}") break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` ```javascript title="JavaScript SDK" highlight={5} expandable import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "", baseUrl: "https://api.eu.assemblyai.com", }); // const audioFile = './local_file.mp3' const audioFile = "https://assembly.ai/wildfires.mp3"; const params = { audio: audioFile, speech_models: ["universal-3-5-pro", "universal-2"], language_detection: true, }; const run = async () => { const transcript = await client.transcripts.transcribe(params); console.log(transcript.text); }; run(); ``` ```javascript title="JavaScript" highlight={4} expandable import fs from "fs-extra"; const baseUrl = "https://api.eu.assemblyai.com"; const headers = { authorization: "", }; const path = "./my-audio.mp3"; const audioData = await fs.readFile(path); let res = await fetch(`${baseUrl}/v2/upload`, { method: "POST", headers, body: audioData, }); if (!res.ok) throw new Error(`Error: ${res.status}`); const uploadResponse = await res.json(); const uploadUrl = uploadResponse.upload_url; const data = { audio_url: uploadUrl, // You can also use a URL to an audio or video file on the web speech_models: ["universal-3-5-pro", "universal-2"], language_detection: true, }; const url = `${baseUrl}/v2/transcript`; res = await fetch(url, { method: "POST", headers: { ...headers, "Content-Type": "application/json" }, body: JSON.stringify(data), }); if (!res.ok) throw new Error(`Error: ${res.status}`); const response = await res.json(); const transcriptId = response.id; const pollingEndpoint = `${baseUrl}/v2/transcript/${transcriptId}`; while (true) { res = await fetch(pollingEndpoint, { headers }); if (!res.ok) throw new Error(`Error: ${res.status}`); const transcriptionResult = await res.json(); if (transcriptionResult.status === "completed") { console.log(transcriptionResult.text); break; } else if (transcriptionResult.status === "error") { throw new Error(`Transcription failed: ${transcriptionResult.error}`); } else { await new Promise((resolve) => setTimeout(resolve, 3000)); } } ``` --- # Rate limits URL: https://www.assemblyai.com/docs/pre-recorded-audio/rate-limits Source: docs/pre-recorded-audio/rate-limits.mdx Navigation: Pre-recorded STT > Advanced Description: Rate limits documentation. When you submit audio files to the `/v2/transcript` endpoint, AssemblyAI processes them in parallel up to your account's rate limit. The rate limit is the maximum number of transcription jobs that can be actively processing at the same time. ## Default limits | Account type | Rate limit (parallel transcriptions) | | ------------ | ------------------------------------ | | Free | 5 | | Paid | 200+ | **Need a higher rate limit?** Our services are infinitely scalable and we offer custom rate limits that scale to support any workload at no additional cost. If you need a higher rate limit, please either [contact our Sales team](https://www.assemblyai.com/contact) or send an email to [support@assemblyai.com](mailto:support@assemblyai.com). ## What happens when you hit the limit If you submit a transcription that would exceed your rate limit, it is added to a queue. Queued transcriptions are processed automatically in FIFO (first-in, first-out) order as previously submitted transcriptions complete. If you exceed your rate limit, you will receive an email stating that your transcripts have been throttled. You will only receive this email once per day. ![Email notifying user that their transcripts have been throttled due to exceeding the rate limit](/assets/pages/guides/limit.png) If your account balance goes below zero, your rate limit will be reduced to 1. ## HTTP rate limits In addition to the parallel-job rate limit, there is an HTTP rate limit for the API that restricts accounts to a maximum of 20,000 requests per five minutes. This limit counts all requests across endpoints — submissions (POST) and polling (GET) combined — and exceeding it returns an HTTP `403`. To stay under it, use [webhooks](/pre-recorded-audio/webhooks) instead of frequent polling, or widen and jitter your polling interval. See [Polling without exceeding the rate limit](/pre-recorded-audio/guides/bulk-transcription-and-load-tests-at-scale#polling-without-exceeding-the-rate-limit) for a worked example. ## Check your limit You can view your current rate limit on the [Rate Limits page](https://www.assemblyai.com/dashboard/home) of your dashboard. With the current version of multi-project support, rate limiting is applied at the account level, not at the project level. This means that the rate limits for each API key mirror the rate limits for the account. For example, if an account has a rate limit of 200, each API key for that account will be able to process up to 200 requests in parallel. ## Related pages - [Real-time STT rate limits](/streaming/rate-limits) - [Account Management](/account-management) - [Transcript status](/pre-recorded-audio/check-transcript-status) --- # Migration guide: Deepgram to AssemblyAI URL: https://www.assemblyai.com/docs/pre-recorded-audio/migration-guides/dg_to_aai Source: docs/pre-recorded-audio/migration-guides/dg_to_aai.mdx Navigation: Pre-recorded STT > Guides > Migration guides Description: Migration guide: Deepgram to AssemblyAI documentation. This guide walks through the process of migrating from Deepgram to AssemblyAI for transcribing pre-recorded audio. ## Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. ## Side-By-Side Code Comparison Below is a side-by-side comparison of a basic snippet to transcribe a **local file** by Deepgram and AssemblyAI: ```python expandable from deepgram import ( DeepgramClient, PrerecordedOptions, FileSource, ) API_KEY = "YOUR_DG_API_KEY" AUDIO_FILE = "./example.wav" def main(): try: deepgram = DeepgramClient(API_KEY) with open(AUDIO_FILE, "rb") as file: buffer_data = file.read() payload: FileSource = { "buffer": buffer_data, } options = PrerecordedOptions( model="nova-2", smart_format=True, diarize=True ) response = deepgram.listen.prerecorded.v("1").transcribe_file(payload, options) print(response.to_json(indent=4)) except Exception as e: print(f"Exception: {e}") if name == "main": main() ``` ```python expandable import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() audio_file = "./example.wav" config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels=True, ) transcript = transcriber.transcribe(audio_file, config) if transcript.status == aai.TranscriptStatus.error: print(f"Transcription failed: {transcript.error}") exit(1) print(transcript.text) for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` Below is a side-by-side comparison of a basic snippet to transcribe a **publicly-accessible URL** by Deepgram and AssemblyAI: ```python expandable from deepgram import ( DeepgramClient, PrerecordedOptions ) API_KEY = "YOUR_DG_API_KEY" AUDIO_URL = { "url": "https://dpgr.am/spacewalk.wav" } def main(): try: deepgram = DeepgramClient(API_KEY) options = PrerecordedOptions( model="nova-2", smart_format=True, diarize=True ) response = deepgram.listen.prerecorded.v("1").transcribe_url(AUDIO_URL, options) print(response.to_json(indent=4)) except Exception as e: print(f"Exception: {e}") if name == "main": main() ``` ```python expandable import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() audio_file = ( "https://assembly.ai/sports_injuries.mp3" ) config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels=True, ) transcript = transcriber.transcribe(audio_file, config) if transcript.status == aai.TranscriptStatus.error: print(f"Transcription failed: {transcript.error}") exit(1) print(transcript.text) for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` Here are helpful things to know about our `transcribe` method: - The SDK handles polling under the hood - Transcript is directly accessible via `transcript.text` - English is the default language. We recommend specifying `speech_models=["universal-3-5-pro", "universal-2"]` for the highest accuracy - We have a [cookbook for error handling common errors](/pre-recorded-audio/guides/common_errors_and_solutions) when using our API. ## Installation ```python from deepgram import ( DeepgramClient, PrerecordedOptions, FileSource, ) API_KEY = "YOUR_DG_API_KEY" deepgram = DeepgramClient(API_KEY) ``` ```python import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() ``` When migrating from Deepgram to AssemblyAI, you'll first need to handle authentication and SDK setup: Get your API key from your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home) \ To follow this guide, install AssemblyAI's Python SDK by typing this code into your terminal: \ `pip install assemblyai` Things to know: - Store your API key securely in an environment variable - API key authentication works the same across all AssemblyAI SDKs ## Audio File Sources ```python expandable # Local Files AUDIO_FILE = "example.wav" with open(AUDIO_FILE, "rb") as file: buffer_data = file.read() payload: FileSource = { "buffer": buffer_data, } options = PrerecordedOptions( smart_format=True, summarize="v2", ) file_response = deepgram.listen.rest.v("1").transcribe_file(payload, options) json = file_response.to_json() #Public URLs AUDIO_URL = { "url": "https://static.deepgram.com/examples/Bueller-Life-moves-pretty-fast.wav" } options = PrerecordedOptions( smart_format=True, summarize="v2" ) url_response = deepgram.listen.rest.v("1").transcribe_url(AUDIO_URL, options) json = url_response.to_json() ``` ```python transcriber = aai.Transcriber() # Local Files transcript = transcriber.transcribe("./audio.mp3") # Public URLs transcript = transcriber.transcribe("https://example.com/audio.mp3") # S3 files (using pre-signed URLs) s3_client = boto3.client('s3') presigned_url = s3_client.generate_presigned_url( 'get_object', Params={'Bucket': 'my-bucket', 'Key': 'audio.mp3'}, ExpiresIn=3600 ) transcript = transcriber.transcribe(presigned_url) ``` Here are helpful things to know when migrating your audio input handling: - There's no need to specify the audio format to AssemblyAI - it's auto-detected. AssemblyAI accepts almost every audio/video file type: [here is a full list of all our supported file types](/faq/what-audio-and-video-file-types-are-supported-by-your-api) - Our SDK handles file upload and transcription automatically in one step - For S3 files, you'll need to generate pre-signed URLs ([see example in cookbook](/pre-recorded-audio/guides/transcribe_from_s3)) ## Adding Features ```python options = PrerecordedOptions( model="nova-2", smart_format=True, diarize=True, detect_entities=True ) response = deepgram.listen.prerecorded.v("1").transcribe_url(AUDIO_URL, options) ``` ```python config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels=True, # Speaker diarization auto_chapters=True, # Auto chapter detection entity_detection=True, # Named entity detection ) transcript = transcriber.transcribe(audio_url, config) # Access speaker labels for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` Key differences: - Use `aai.TranscriptionConfig` to specify any extra features that you wish to use - The results for Speaker Diarization are stored in `transcript.utterances`. To see the full transcript response object, refer to our [API Reference](/api-reference). - Check our [documentation](/speech-understanding/getting-started) for our full list of available features and their parameters --- # Migration guide: OpenAI to AssemblyAI URL: https://www.assemblyai.com/docs/pre-recorded-audio/migration-guides/oai_to_aai Source: docs/pre-recorded-audio/migration-guides/oai_to_aai.mdx Navigation: Pre-recorded STT > Guides > Migration guides Description: Migration guide: OpenAI to AssemblyAI documentation. This guide walks through the process of migrating from OpenAI to AssemblyAI for transcribing pre-recorded audio. ## Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. ## Side-By-Side Code Comparison Below is a side-by-side comparison of a basic snippet to transcribe a **local file** by OpenAI and AssemblyAI: ```python from openai import OpenAI api_key = "YOUR_OPENAI_API_KEY" client = OpenAI(api_key) audio_file = open("./example.wav", "rb") transcript = client.audio.transcriptions.create( model = "whisper-1", file = audio_file ) print(transcript.text) ``` ```python import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() audio_file = "./example.wav" config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, ) transcript = transcriber.transcribe(audio_file, config) if transcript.status == aai.TranscriptStatus.error: print(f"Transcription failed: {transcript.error}") exit(1) print(transcript.text) ``` Here are helpful things to know about our `transcribe` method: - The SDK handles polling under the hood - Transcript is directly accessible via `transcript.text` - English is the default language. We recommend specifying `speech_models=["universal-3-5-pro", "universal-2"]` for the highest accuracy - We have a [cookbook for error handling common errors](/pre-recorded-audio/guides/common_errors_and_solutions) when using our API. ## Installation ```python from openai import OpenAI api_key = "YOUR_OPENAI_API_KEY" client = OpenAI(api_key) ``` ```python import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() ``` When migrating from OpenAI to AssemblyAI, you'll first need to handle authentication and SDK setup: Get your API key from your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home) \ To follow this guide, install AssemblyAI's Python SDK by typing this code into your terminal: \ `pip install assemblyai` Things to know: - Store your API key securely in an environment variable - API key authentication works the same across all AssemblyAI SDKs ## Audio File Sources ```python client = OpenAI() # Local Files audio_file = open("./example.wav", "rb") transcript = client.audio.transcriptions.create( model = "whisper-1", file = audio_file ) ``` ```python transcriber = aai.Transcriber() # Local Files transcript = transcriber.transcribe("./audio.mp3") # Public URLs transcript = transcriber.transcribe("https://example.com/audio.mp3") ``` Here are helpful things to know when migrating your audio input handling: - AssemblyAI natively supports transcribing publicly accessible audio URLs (for example, S3 URLs), the Whisper API only natively supports transcribing local files. - There's no need to specify the audio format to AssemblyAI - it's auto-detected. AssemblyAI accepts almost every audio/video file type: [here is a full list of all our supported file types](/faq/what-audio-and-video-file-types-are-supported-by-your-api) - The Whisper API only supports file sizes up to 25MB, AssemblyAI supports file sizes up to 5GB. ## Adding Features ```python transcript = client.audio.transcriptions.create( file = audio_file, prompt = "INSERT_PROMPT", # Optional text to guide the model's style language = "en", # Set language code model = "whisper-1", response_format = "verbose_json", timestamp_granularities = ["word"] ) # Access word-level timestamps print(transcript.words) ``` ```python config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels = True, # Speaker diarization sentiment_analysis=True, # Sentiment Analysis entity_detection = True, # Named entity detection ) transcript = transcriber.transcribe(audio_url, config) # Access word-level timestamps print(transcript.words) # Access speaker labels for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` Key differences: - OpenAI does not offer speech understanding features for their speech-to-text API - Use `aai.TranscriptionConfig` to specify any extra features that you wish to use - With AssemblyAI, timestamp granularity is word-level by default - The results for Speaker Diarization are stored in `transcript.utterances`. To see the full transcript response object, refer to our [API Reference](/api-reference). - Check our [documentation](/speech-understanding/getting-started) for our full list of available features and their parameters - If you want to send a custom prompt to an LLM, you can use [LLM Gateway](/llm-gateway/quickstart) to apply the model to your transcribed audio files. --- # Migration guide: AWS Transcribe to AssemblyAI URL: https://www.assemblyai.com/docs/pre-recorded-audio/migration-guides/aws_to_aai Source: docs/pre-recorded-audio/migration-guides/aws_to_aai.mdx Navigation: Pre-recorded STT > Guides > Migration guides Description: Migration guide: AWS Transcribe to AssemblyAI documentation. This guide walks through the process of migrating from AWS Transcribe to AssemblyAI for transcribing pre-recorded audio. ## Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. ## Side-by-side code comparison Below is a side-by-side comparison of a basic snippet to transcribe a file by AWS Transcribe and AssemblyAI: ```python expandable import time import boto3 def transcribe_file(job_name, file_uri, transcribe_client): transcribe_client.start_transcription_job( TranscriptionJobName=job_name, Media={"MediaFileUri": file_uri}, MediaFormat="wav", LanguageCode="en-US", ) max_tries = 60 while max_tries > 0: max_tries -= 1 job = transcribe_client.get_transcription_job( TranscriptionJobName=job_name ) job_status = job["TranscriptionJob"]["TranscriptionJobStatus"] if job_status in ["COMPLETED", "FAILED"]: print(f"Job {job_name} is {job_status}.") if job_status == "COMPLETED": print( f"Download the transcript from\n" f"\t{job['TranscriptionJob']['Transcript']['TranscriptFileUri']}." ) break else: print(f"Waiting for {job_name}. Current status is {job_status}.") time.sleep(10) def main(): transcribe_client = boto3.client("transcribe") file_uri = "s3://test-transcribe/answer2.wav" transcribe_file("Example-job", file_uri, transcribe_client) if name == "main": main() ``` ```python expandable import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() # You can use a local filepath: # audio_file = "./example.mp3" # Or use a publicly-accessible URL: audio_file = ( "https://assembly.ai/sports_injuries.mp3" ) config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels=True, ) transcript = transcriber.transcribe(audio_file, config) if transcript.status == aai.TranscriptStatus.error: print(f"Transcription failed: {transcript.error}") exit(1) print(transcript.text) for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` ## Installation ```python import boto3 import time transcribe_client = boto3.client("transcribe") ``` ```python import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() ``` When migrating from AWS to AssemblyAI, you'll first need to handle authentication and SDK setup: Get your API key from your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home) Things to know: - Store your API key securely in an environment variable - API key authentication works the same across all AssemblyAI SDKs ## Audio File Sources ```python def transcribe_file(job_name, file_uri, transcribe_client): transcribe_client.start_transcription_job( TranscriptionJobName=job_name, Media={"MediaFileUri": file_uri}, MediaFormat="wav", LanguageCode="en-US", ) ``` ```python transcriber = aai.Transcriber() # Local files transcript = transcriber.transcribe("./audio.mp3") # Public URLs transcript = transcriber.transcribe("https://example.com/audio.mp3") # S3 files (using pre-signed URLs) s3_client = boto3.client('s3') presigned_url = s3_client.generate_presigned_url( 'get_object', Params={'Bucket': 'my-bucket', 'Key': 'audio.mp3'}, ExpiresIn=3600 ) transcript = transcriber.transcribe(presigned_url) ``` Here are helpful things to know when migrating your audio input handling: - There's no need to specify the audio format to AssemblyAI - it's auto-detected. AssemblyAI accepts almost every audio/video file type: [here is a full list of all our supported file types](/faq/what-audio-and-video-file-types-are-supported-by-your-api) - Our SDK handles file upload and transcription automatically in one step - For S3 files, you'll need to generate pre-signed URLs ([see example in cookbook](/pre-recorded-audio/guides/transcribe_from_s3)) ## Basic Transcription ```python while max_tries > 0: max_tries -= 1 job = transcribe_client.get_transcription_job( TranscriptionJobName=job_name ) job_status = job["TranscriptionJob"]["TranscriptionJobStatus"] if job_status in ["COMPLETED", "FAILED"]: break time.sleep(10) ``` ```python transcriber = aai.Transcriber() # Local files transcript = transcriber.transcribe("./audio.mp3") # Public URLs transcript = transcriber.transcribe("https://example.com/audio.mp3") # S3 files (using pre-signed URLs) s3_client = boto3.client('s3') presigned_url = s3_client.generate_presigned_url( 'get_object', Params={'Bucket': 'my-bucket', 'Key': 'audio.mp3'}, ExpiresIn=3600 ) transcript = transcriber.transcribe(presigned_url) ``` Here are helpful things to know about our `transcribe` method: - The SDK handles polling under the hood - Transcript is directly accessible via `transcript.text` - English is the default language. We recommend specifying `speech_models=["universal-3-5-pro", "universal-2"]` for the highest accuracy - We have a [cookbook for error handling common errors](/pre-recorded-audio/guides/common_errors_and_solutions) when using our API. ## Adding Features ```python transcribe_client.start_transcription_job( TranscriptionJobName=job_name, Media={"MediaFileUri": file_uri}, Settings={ "ShowSpeakerLabels": True, "MaxSpeakerLabels": 2 } ) ``` ```python config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels=True, # Speaker diarization auto_chapters=True, # Auto chapter detection entity_detection=True, # Named entity detection ) transcript = transcriber.transcribe(audio_file, config) # Access speaker labels for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` Key differences: - Use `aai.TranscriptionConfig` to specify any extra features that you wish to use - The results for Speaker Diarization are stored in `transcript.utterances`. To see the full transcript response object, refer to our [API Reference](/api-reference). - Check our [documentation](/speech-understanding/getting-started) for our full list of available features and their parameters --- # Migration guide: Google Speech-to-Text to AssemblyAI URL: https://www.assemblyai.com/docs/pre-recorded-audio/migration-guides/google_to_aai Source: docs/pre-recorded-audio/migration-guides/google_to_aai.mdx Navigation: Pre-recorded STT > Guides > Migration guides Description: Migration guide: Google Speech-to-Text to AssemblyAI documentation. This guide walks through the process of migrating from Google Speech-to-Text (STT) to AssemblyAI for transcribing pre-recorded audio. ## Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. ## Side-by-side code comparison Below is a side-by-side comparison of a basic snippet to transcribe a file by Google Speech-to-Text and AssemblyAI. ```python expandable from google.cloud import speech client = speech.SpeechClient() audio = speech.RecognitionAudio( uri="gs://cloud-samples-tests/speech/Google_Gnome.wav" ) config = speech.RecognitionConfig( encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=16000, language_code="en-US", model="video", # Chosen model ) operation = client.long_running_recognize(config=config, audio=audio) print("Waiting for operation to complete...") response = operation.result(timeout=90) for i, result in enumerate(response.results): alternative = result.alternatives[0] print("-" * 20) print(f"First alternative of result {i}") print(f"Transcript: {alternative.transcript}") ``` ```python expandable import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() # You can use a local filepath: # audio_file = "./example.mp3" # Or use a publicly-accessible URL: audio_file = ( "https://assembly.ai/sports_injuries.mp3" ) config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, ) transcript = transcriber.transcribe(audio_file, config) if transcript.status == aai.TranscriptStatus.error: print(f"Transcription failed: {transcript.error}") exit(1) print(transcript.text) ``` ## Installation ```python from google.cloud import speech client = speech.SpeechClient() ``` ```python import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" transcriber = aai.Transcriber() ``` When migrating from Google Speech-to-Text to AssemblyAI, you'll first need to handle authentication and SDK setup: Get your API key from your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home). Things to know: - Store your API key securely in an environment variable - API key authentication works the same across all AssemblyAI SDKs ## Audio File Sources ```python audio = speech.RecognitionAudio(uri="gs://cloud-samples-tests/speech/Google_Gnome.wav") config = speech.RecognitionConfig( encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=16000, language_code="en-US", model="video", # Chosen model ) operation = client.long_running_recognize(config=config, audio=audio) ``` ```python transcriber = aai.Transcriber() # Local files transcript = transcriber.transcribe("./audio.mp3") # Public URLs transcript = transcriber.transcribe("https://example.com/audio.mp3") # S3 files (using pre-signed URLs) s3_client = boto3.client('s3') presigned_url = s3_client.generate_presigned_url( 'get_object', Params={'Bucket': 'my-bucket', 'Key': 'audio.mp3'}, ExpiresIn=3600 ) transcript = transcriber.transcribe(presigned_url) ``` Here are helpful things to know when migrating your audio input handling: - There's no need to specify the audio encoding format when using AssemblyAI - we have a transcoding pipeline under the hood which works on all [supported file types](/faq/what-audio-and-video-file-types-are-supported-by-your-api) so that you can get the most accurate transcription. - You can submit a local file, URL, stream, buffer, blob, etc., directly to our transcriber. Check out some common ways you can host audio files [here](/pre-recorded-audio/guides/transcribe_from_s3). - You can transcribe audio files that are up to 10 hours long and you can transcribe multiple files in parallel. The default amount of jobs you can transcribe at once is 200 while on the PAYG plan. ## Basic Transcription ```python print("Waiting for operation to complete...") response = operation.result(timeout=90) for i, result in enumerate(response.results): alternative = result.alternatives[0] print("-" * 20) print(f"First alternative of result {i}") print(f"Transcript: {alternative.transcript}") ``` ```python transcriber = aai.Transcriber() config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, ) transcript = transcriber.transcribe(audio_file, config) if transcript.status == aai.TranscriptStatus.error: print(f"Transcription failed: {transcript.error}") else: print(transcript.text) ``` Here are helpful things to know about our `transcribe` method: - The SDK handles polling under the hood. - The full transcript is directly accessible via `transcript.text`. - English is the default language. We recommend specifying `speech_models=["universal-3-5-pro", "universal-2"]` for the highest accuracy. - We have a [cookbook for error handling common errors](/pre-recorded-audio/guides/common_errors_and_solutions) when using our API. ## Adding Features ```python config = speech.RecognitionConfig( encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=8000, language_code="en-US", enable_speaker_diarization=True, # Speaker diarization diarization_speaker_count=2, # Specify amount of speakers profanity_filter=True # Remove profanity from transcript ) ``` ```python config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels=True, # Speaker diarization filter_profanity=True, # Remove profanity from transcript speakers_expected=2, # Specify amount of speakers in audio ) transcript = transcriber.transcribe(audio_file, config) # Access speaker labels for utterance in transcript.utterances: print(f"Speaker {utterance.speaker}: {utterance.text}") ``` Key differences: - Use `aai.TranscriptionConfig` to specify any extra features that you wish to use. - The results for Speaker Diarization are stored in `transcript.utterances`. To see the full transcript response object, refer to our [API Reference](/api-reference). - Check our [documentation](/speech-understanding/getting-started) for our full list of available features and their parameters. --- # Migration guide: Gladia to AssemblyAI URL: https://www.assemblyai.com/docs/pre-recorded-audio/migration-guides/gladia_to_aai Source: docs/pre-recorded-audio/migration-guides/gladia_to_aai.mdx Navigation: Pre-recorded STT > Guides > Migration guides Description: Migration guide: Gladia to AssemblyAI documentation. This guide walks through the process of migrating from Gladia to AssemblyAI for transcribing pre-recorded audio. ## Get started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your [dashboard](https://www.assemblyai.com/dashboard/home). If you'd prefer to use one of our official SDKs, check our [documentation for the full list of available SDKs](/pre-recorded-audio/guides/do-more-with-sdk). The [Gladia documentation](https://docs.gladia.io/chapters/pre-recorded-stt/getting-started) uses cURL commands to demonstrate API usage. In this guide, we will use Python code snippets to illustrate the same functionality across both APIs. If you prefer to use cURL, you can find the equivalent commands in the [AssemblyAI API Reference](/api-reference/overview). ## Side-by-side code comparison Below is a side-by-side comparison of a basic snippet to transcribe pre-recorded audio with Gladia and AssemblyAI: ```python expandable import requests import time base_url = "https://api.gladia.io" headers = { "x-gladia-key": "" } with open("./my-audio.mp3", "rb") as f: files = {"audio": ("my-audio.mp3", f, "audio/mp3")} response = requests.post(base_url + "/v2/upload", headers=headers, files=files) upload_url = response.json()["audio_url"] data = { "audio_url": upload_url # You can also use a URL to an audio or video file on the web. } url = base_url + "/v2/pre-recorded" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] # You can also use response.json()['result_url'] to get the polling_endpoint directly. polling_endpoint = url + "/" + transcript_id while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript['status'] == 'done': print(f"Full Transcript: {transcript['result']['transcription']['full_transcript']}") break elif transcript['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcript['error_code']}") else: time.sleep(3) ``` ```python expandable import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, # You can also use a URL to an audio or video file on the web "speech_models": ["universal-3-5-pro", "universal-2"], "language_detection": True, } url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) transcript_id = response.json()['id'] polling_endpoint = url + "/" + transcript_id while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript['status'] == 'completed': print(f"Full Transcript: {transcript['text']}") break elif transcript['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) ``` ## Installation and authentication ```python import requests import time base_url = "https://api.gladia.io" headers = { "x-gladia-key": "" } ``` ```python import requests import time base_url = "https://api.assemblyai.com" headers = { "authorization": "" } ``` When migrating from Gladia to AssemblyAI, you'll first need to handle authentication: Get your API key from your [AssemblyAI dashboard](https://www.assemblyai.com/dashboard/home). Things to know: - Store your API key securely in an environment variable. - We support the ability to create [multiple API keys and projects](/account-management) to help you track and manage seperate environments. Gladia uses the `x-gladia-key` HTTP header for authentication, while AssemblyAI uses the `authorization` header. ## Audio file sources You can provide either a locally stored audio file or a publicly accessible URL. ```python # Local Files with open("./my-audio.mp3", "rb") as f: files = {"audio": ("my-audio.mp3", f, "audio/mp3")} response = requests.post(base_url + "/v2/upload", headers=headers, files=files) upload_url = response.json()["audio_url"] data = { "audio_url": upload_url } #Public URLs audio_file = "https://assembly.ai/sports_injuries.mp3" data = { "audio_url": audio_file } ``` ```python expandable # Local Files with open("./my-audio.mp3", "rb") as f: response = requests.post(base_url + "/v2/upload", headers=headers, data=f) upload_url = response.json()["upload_url"] data = { "audio_url": upload_url, "speech_models": ["universal-3-5-pro", "universal-2"], "language_detection": True, } # Public URLs audio_file = "https://assembly.ai/sports_injuries.mp3" data = { "audio_url": audio_file, "speech_models": ["universal-3-5-pro", "universal-2"], "language_detection": True, } ``` ## Basic transcription and polling the transcription status Make a `POST` request to the [/v2/pre-recorded](https://docs.gladia.io/docs/api-reference/v2/pre-recorded/init) endpoint. ```python url = base_url + "/v2/pre-recorded" response = requests.post(url, json=data, headers=headers) ``` Every few seconds, make a `GET` request to the [/v2/pre-recorded/:transcript_id](https://docs.gladia.io/docs/api-reference/v2/pre-recorded/get) endpoint until the transcription status is `'done'`. ```python transcript_id = response.json()['id'] # You can also use response.json()['result_url'] to get the polling_endpoint directly. polling_endpoint = url + "/" + transcript_id while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript['status'] == 'done': print(f"Full Transcript: {transcript['result']['transcription']['full_transcript']}") break elif transcript['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcript['error_code']}") else: time.sleep(3) ``` Make a `POST` request to the [/v2/transcript](/api-reference/transcripts/submit) endpoint. ```python url = base_url + "/v2/transcript" response = requests.post(url, json=data, headers=headers) ``` Every few seconds, make a `GET` request to the [/v2/transcript/:transcript_id](/api-reference/transcripts/get) endpoint until the transcription status is `'completed'`. ```python transcript_id = response.json()['id'] polling_endpoint = url + "/" + transcript_id while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript['status'] == 'completed': print(f"Full Transcript: {transcript['text']}") break elif transcript['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) ``` **Transcription status** Note that our APIs possible values for transcription status are `queued`, `processing`, `completed`, and `error`. Check out the [AssemblyAI API Reference](/api-reference/transcripts/get#response.body.status) for the full list of possible transcription status values. If you'd rather not poll the API, you can use our [SDKs](/pre-recorded-audio/guides/do-more-with-sdk) which handle polling internally. Alternatively, you can also use [webhooks](/pre-recorded-audio/webhooks) to get notified when your transcript is complete. Here are helpful things to know when migrating your audio input handling: - Both AssemblyAI and Gladia allow you the option of uploading a local file or specifying a publicly accessible URL. - There's no need to specify the audio format to AssemblyAI - it's auto-detected. AssemblyAI accepts almost every audio/video file type: [here is a full list of all our supported file types](/faq/what-audio-and-video-file-types-are-supported-by-your-api) - For self-hosted or pre-signed URLs (i.e. S3), [see our example in this cookbook](/pre-recorded-audio/guides/transcribe_from_s3). ## Adding features ```python data = { "audio_url": upload_url, "diarization": True, # Speaker diarization "chapterization": True, # Auto chapter detection "named_entity_recognition": True # Named entity detection } # Access speaker labels for utterance in transcript['result']['transcription']['utterances']: print(f"Speaker {utterance['speaker']}: {utterance['text']}") # Access auto chapters for chapter in transcript['result']['chapterization']['results']: print(f"{chapter['start']} - {chapter['end']}: {chapter['headline']}") # Access named entities for entity in transcript['result']['named_entity_recognition']['results']: print(entity['text']) print(entity['entity_type']) print(f"Timestamp: {entity['start']} - {entity['end']}\n") ``` ```python expandable data = { "audio_url": upload_url, "speech_models": ["universal-3-5-pro", "universal-2"], "language_detection": True, "speaker_labels": True, # Speaker diarization "auto_chapters": True, # Auto chapter detection "entity_detection": True # Named entity detection } # Access speaker labels for utterance in transcript['utterances']: print(f"Speaker {utterance['speaker']}: {utterance['text']}") # Access auto chapters for chapter in transcript['chapters']: print(f"{chapter['start']} - {chapter['end']}: {chapter['headline']}") # Access named entities for entity in transcript['entities']: print(entity['text']) print(entity['entity_type']) print(f"Timestamp: {entity['start']} - {entity['end']}\n") ``` Key differences: - Make sure to note any differences in parameters or response structure. If using [Speaker Diarization](/pre-recorded-audio/label-speakers), for example: - Parameters: AssemblyAI uses `speaker_labels`, while Gladia uses `diarization`. - Response: AssemblyAI uses `transcript.utterances`, while Gladia uses `transcript.result.transcription.utterances`. - Make sure to review each API reference for the full list of parameters and response objects. - [AssemblyAI API Reference](/api-reference/transcripts/submit) - [Gladia API Reference](https://docs.gladia.io/docs/api-reference/v2/pre-recorded/init) --- # Best Practices for building Meeting Notetakers URL: https://www.assemblyai.com/docs/meeting-notetaker-best-practices Source: docs/meeting-notetaker-best-practices.mdx Navigation: Pre-recorded STT > Guides > Tutorials Description: Complete guide for building meeting notetakers with AssemblyAI ## Introduction Building a robust meeting notetaker requires careful consideration of accuracy, latency, speaker identification, and real-time capabilities. This guide addresses common questions and provides practical solutions for both post-call and live meeting transcription scenarios. ## Why AssemblyAI for Meeting Notetakers? AssemblyAI stands out as the premier choice for meeting notetakers with several key advantages: ### Industry-Leading Accuracy with Pre-recorded Audio - **93.3%+ transcription accuracy** ensures reliable meeting documentation - **2.9% speaker diarization error rate** for precise "who said what" attribution - **Speech Understanding** integration for intelligent post-processing and insights - **Keyterms prompt** allows providing meeting context to improve accuracy of transcription ### Streaming with Universal-3.5 Pro As meeting notetakers evolve toward real-time capabilities, AssemblyAI's Universal-3.5 Pro Streaming model (`universal-3-5-pro`) offers significant benefits: - **Speaker diarization** available for both pre-recorded and streaming transcription - **Ultra-low latency (~300ms)** enables live transcription without delays - **Format turns** feature provides structured, readable output in real-time - **Keyterms prompt** allows providing meeting context to improve accuracy of transcription ### End-to-End Voice AI Platform Unlike fragmented solutions, AssemblyAI provides a unified API for: - Transcription with speaker diarization - Automatic language detection and code switching - Boosting accuracy via meeting context with keyterms prompt - Speech Understanding tasks like speaker identification, translation, and transcript styling - Post-processing workflows with custom prompting - from summarization to completely custom workflows - Real-time and batch processing of pre-recorded audio in a single platform ## When Should I Use Pre-recorded vs Streaming for Meeting Notetakers? Understanding when to use pre-recorded versus streaming speech-to-text is critical for building the right meeting notetaker. ### Pre-recorded Speech-to-text **Post-call analysis** - Meeting already happened, you have the full recording - **Highest accuracy needed** - Pre-recorded models have higher accuracy (93.3%+) - **Speaker diarization is critical** - Pre-recorded has 2.9% speaker error rate - **Broad language support** - Need any of 99+ languages - **Advanced features required** - Summarization, sentiment analysis, entity detection, PII redaction, speaker identification - **Batch processing** - Processing multiple recordings at once - **Quality over speed** - Can wait seconds/minutes for perfect results **Best for:** Zoom/Teams/Meet recording uploads, compliance, documentation, post-call summaries, searchable archives ### Streaming Speech-to-text **Live meetings** - Transcribing as the meeting happens You should use streaming when you need to display a live transcript of text to users as they are speaking. With Universal-3.5 Pro Streaming, accuracy is closer to pre-recorded, but pre-recorded will always be the most accurate option. - **Real-time captions** - Displaying subtitles/captions to participants during calls - **Immediate feedback** - Need transcription within ~300ms - **Interactive features** - Live note-taking, real-time keyword detection, action item alerts - **No recording available** - Processing live audio only **Best for:** Live captions, real-time note-taking apps, accessibility features, live keyword alerts **Streaming is billed per session** Streaming is billed on the total duration that your WebSocket connection stays open, not on the amount of audio you send. For long-running meetings, make sure to terminate sessions when the meeting ends to avoid being billed for idle time. See [Billing and pricing](/billing-and-pricing) for details. ### Hybrid Approach (Recommended) Many successful meeting notetakers use **both** pre-recorded and streaming speech-to-text: 1. **Streaming during the call** - Provide live captions and real-time notes to participants 2. **Pre-recorded after the call** - Generate high-quality transcript with speaker labels, summary, and insights This gives users immediate value during meetings while providing comprehensive documentation afterward. **Example workflow:** - User joins meeting → Start streaming for live captions - Meeting ends → Upload recording to pre-recorded API for final transcript with speaker names - Generate meeting summary, action items, and searchable archive from pre-recorded transcript ## What Languages and Features for a Meeting Notetaker? ### Pre-Recorded Meetings For post-call analysis, AssemblyAI supports: **Languages**: - 99 languages supported - Automatic Language Detection to route to the most spoken language - Code Switching to preserve changes in speech between languages **Core Features**: - Speaker diarization (1-10 speakers by default, expandable to any min/max) - Multichannel audio support (each channel = one speaker) - Automatic formatting, punctuation, and capitalization - Keyterms prompting for boosting domain-specific terms **Speech Understanding Models**: - Summarization for meeting recaps - Sentiment analysis for meeting tone assessment - Entity detection for extracting key information - Speaker identification to map generic labels to actual names/roles - Translation between 86 languages ### Real-Time Streaming For live meeting transcription: **Languages**: - English-only model (default) - Multilingual model supporting English, Spanish, French, German, Portuguese, and Italian ### Streaming (Universal-3.5 Pro Streaming) - Speaker diarization for identifying who is speaking - Partial and final transcripts for responsive UI - Format turns for structured, readable output - Keyterms prompt for contextual accuracy See the [Universal-3.5 Pro Streaming documentation](/streaming/getting-started/transcribe-streaming-audio) for full details. ## How Can I Get Started Building a Post-Call Meeting Notetaker? Here's a complete example implementing pre-recorded transcription with all essential features: ```python expandable import assemblyai as aai import asyncio from typing import Dict, List from assemblyai.types import ( SpeakerOptions, LanguageDetectionOptions, PIIRedactionPolicy, PIISubstitutionPolicy, ) # Configure API key aai.settings.api_key = "your_api_key_here" async def transcribe_meeting_async(audio_source: str) -> Dict: """ Asynchronously transcribe a meeting recording with full features Args: audio_source: Either a local file path or publicly accessible URL """ # Configure comprehensive meeting analysis config = aai.TranscriptionConfig( # Speaker diarization speaker_labels=True, speakers_expected=None, # Use if you know exact number from Zoom/Meet/Teams speaker_options=SpeakerOptions( min_speakers_expected=2, max_speakers_expected=10 # Set a bit higher than expected; too high can cause over-splitting ), multichannel=False, # Set to True if audio has separate channel per speaker # Language detection language_detection=True, # Auto-detect the most used language language_detection_options=LanguageDetectionOptions( code_switching=True, # Preserve language switches code_switching_confidence_threshold=0.5, ), # Punctuation and formatting punctuate=True, format_text=True, # Boost accuracy of meeting-specific vocabulary keyterms_prompt=["quarterly", "KPI", "roadmap", "deliverables"], # Speech Understanding - commonly used models summarization=True, sentiment_analysis=True, entity_detection=True, redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.organization, PIIRedactionPolicy.occupation, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True ) # Create transcriber transcriber = aai.Transcriber() try: # Submit transcription job transcript = await asyncio.to_thread( transcriber.transcribe, audio_source, config=config ) # Check status if transcript.status == aai.TranscriptStatus.error: raise Exception(f"Transcription failed: {transcript.error}") # Process speaker-labeled utterances print("\n=== SPEAKER-LABELED TRANSCRIPT ===\n") for utterance in transcript.utterances: # Format timestamp start_time = utterance.start / 1000 # Convert to seconds end_time = utterance.end / 1000 # Print formatted utterance print(f"[{start_time:.1f}s - {end_time:.1f}s] Speaker {utterance.speaker}:") print(f" {utterance.text}") print(f" Confidence: {utterance.confidence:.2%}\n") # Print summary data print("\n=== MEETING SUMMARY ===\n") print({ "id": transcript.id, "status": transcript.status, "duration": transcript.audio_duration, "speaker_count": len(set(u.speaker for u in transcript.utterances)), "word_count": len(transcript.words) if transcript.words else 0, "detected_language": transcript.language_code if hasattr(transcript, 'language_code') else None, "summary": transcript.summary, }) return { "transcript": transcript, "utterances": transcript.utterances, "summary": transcript.summary, } except Exception as e: print(f"Error during transcription: {e}") raise async def main(): """ Example usage with error handling """ # Use either local file OR URL (not both) audio_source = "https://assembly.ai/wildfires.mp3" # Or "path/to/recording.mp3" try: result = await transcribe_meeting_async(audio_source) # Additional processing print(f"\nTotal speakers identified: {len(set(u.speaker for u in result['utterances']))}") print(f"Meeting duration: {result['transcript'].audio_duration} seconds") except Exception as e: print(f"Failed to process meeting: {e}") if __name__ == "__main__": asyncio.run(main()) ``` ## How Can I Get Started Building a During-Call Live Meeting Notetaker? Here's a complete example for real-time streaming transcription with meeting-optimized settings: ```python expandable # pip install pyaudio websocket-client import pyaudio import websocket import json import threading import time from urllib.parse import urlencode from datetime import datetime # --- Configuration --- YOUR_API_KEY = "your_api_key" # Keyterms to improve recognition accuracy KEYTERMS = [ "Alice Johnson", "Bob Smith", "Carol Davis", "quarterly review", "action items", "follow up", "deadline", "budget" ] # MEETING NOTETAKER CONFIGURATION (different from voice agents!) CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "format_turns": True, # ALWAYS TRUE for meetings - users need readable text # Meeting-optimized turn detection (wait longer than voice agents) # universal-3-5-pro defaults: min_turn_silence=100ms, max_turn_silence=1000ms "min_turn_silence": 560, # Wait longer for natural pauses (voice agents use ~100ms) "max_turn_silence": 2000, # Allow thinking pauses # Keyterms for accuracy - pass each term as a separate query parameter "keyterms_prompt": KEYTERMS, } API_ENDPOINT_BASE_URL = "wss://streaming.assemblyai.com/v3/ws" API_ENDPOINT = f"{API_ENDPOINT_BASE_URL}?{urlencode(CONNECTION_PARAMS, doseq=True)}" # Audio Configuration FRAMES_PER_BUFFER = 800 # 50ms of audio SAMPLE_RATE = CONNECTION_PARAMS["sample_rate"] CHANNELS = 1 FORMAT = pyaudio.paInt16 # Global variables audio = None stream = None ws_app = None audio_thread = None stop_event = threading.Event() transcript_buffer = [] def on_open(ws): """Called when the WebSocket connection is established.""" print("=" * 80) print(f"[{datetime.now().strftime('%H:%M:%S')}] Meeting transcription started") print(f"Connected to: {API_ENDPOINT_BASE_URL}") print(f"Keyterms configured: {', '.join(KEYTERMS)}") print("=" * 80) print("\nSpeak into your microphone. Press Ctrl+C to stop.\n") def stream_audio(): """Stream audio from microphone to WebSocket""" global stream while not stop_event.is_set(): try: audio_data = stream.read(FRAMES_PER_BUFFER, exception_on_overflow=False) ws.send(audio_data, websocket.ABNF.OPCODE_BINARY) except Exception as e: if not stop_event.is_set(): print(f"Error streaming audio: {e}") break global audio_thread audio_thread = threading.Thread(target=stream_audio) audio_thread.daemon = True audio_thread.start() def on_message(ws, message): """Handle incoming messages from AssemblyAI""" try: data = json.loads(message) msg_type = data.get("type") # Uncomment to see full JSON for debugging: # print("=" * 80) # print(json.dumps(data, indent=2, ensure_ascii=False)) # print("=" * 80) # print() if msg_type == "Begin": session_id = data.get("id", "N/A") print(f"[SESSION] Started - ID: {session_id}\n") elif msg_type == "Turn": end_of_turn = data.get("end_of_turn", False) transcript = data.get("transcript", "") turn_order = data.get("turn_order", 0) end_of_turn_confidence = data.get("end_of_turn_confidence", 0.0) # FOR MEETING NOTETAKERS: Show partials for responsive UI if not end_of_turn and transcript: print(f"\r[LIVE] {transcript}", end="", flush=True) # FOR MEETING NOTETAKERS: Use formatted finals for readable display # (Unlike voice agents which should use utterance for speed) if end_of_turn and transcript: timestamp = datetime.now().strftime('%H:%M:%S') print(f"\n[{timestamp}] {transcript}") print(f" Turn: {turn_order} | Confidence: {end_of_turn_confidence:.2%}") # Detect action items transcript_lower = transcript.lower() if any(term in transcript_lower for term in ["action item", "follow up", "deadline", "assigned to", "todo"]): print(" ⚠️ ACTION ITEM DETECTED!") # Store final transcript transcript_buffer.append({ "timestamp": timestamp, "text": transcript, "turn_order": turn_order, "confidence": end_of_turn_confidence, "type": "final" }) print() elif msg_type == "Termination": audio_duration = data.get("audio_duration_seconds", 0) print(f"\n[SESSION] Terminated - Duration: {audio_duration}s") save_transcript() elif msg_type == "Error": error_msg = data.get("error", "Unknown error") print(f"\n[ERROR] {error_msg}") except json.JSONDecodeError as e: print(f"Error decoding message: {e}") except Exception as e: print(f"Error handling message: {e}") def on_error(ws, error): """Called when a WebSocket error occurs.""" print(f"\n[WEBSOCKET ERROR] {error}") stop_event.set() def on_close(ws, close_status_code, close_msg): """Called when the WebSocket connection is closed.""" print(f"\n[WEBSOCKET] Disconnected - Status: {close_status_code}, Message: {close_msg}") global stream, audio stop_event.set() # Clean up audio stream if stream: if stream.is_active(): stream.stop_stream() stream.close() stream = None if audio: audio.terminate() audio = None if audio_thread and audio_thread.is_alive(): audio_thread.join(timeout=1.0) def save_transcript(): """Save the transcript to a file""" if not transcript_buffer: print("No transcript to save.") return filename = f"meeting_transcript_{datetime.now().strftime('%Y%m%d_%H%M%S')}.txt" with open(filename, "w", encoding="utf-8") as f: f.write("Meeting Transcript\n") f.write(f"Date: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n") f.write(f"Keyterms: {', '.join(KEYTERMS)}\n") f.write("=" * 80 + "\n\n") for entry in transcript_buffer: f.write(f"[{entry['timestamp']}] {entry['text']}\n") f.write(f"Confidence: {entry['confidence']:.2%}\n\n") print(f"Transcript saved to: {filename}") def run(): """Main function to run the streaming transcription""" global audio, stream, ws_app # Initialize PyAudio audio = pyaudio.PyAudio() # Open microphone stream try: stream = audio.open( input=True, frames_per_buffer=FRAMES_PER_BUFFER, channels=CHANNELS, format=FORMAT, rate=SAMPLE_RATE, ) print("Microphone stream opened successfully.") except Exception as e: print(f"Error opening microphone stream: {e}") if audio: audio.terminate() return # Create WebSocketApp ws_app = websocket.WebSocketApp( API_ENDPOINT, header={"Authorization": YOUR_API_KEY}, on_open=on_open, on_message=on_message, on_error=on_error, on_close=on_close, ) # Run WebSocketApp in a separate thread ws_thread = threading.Thread(target=ws_app.run_forever) ws_thread.daemon = True ws_thread.start() try: # Keep main thread alive until interrupted while ws_thread.is_alive(): time.sleep(0.1) except KeyboardInterrupt: print("\n\nCtrl+C received. Stopping transcription...") stop_event.set() # Send termination message to the server if ws_app and ws_app.sock and ws_app.sock.connected: try: terminate_message = {"type": "Terminate"} ws_app.send(json.dumps(terminate_message)) time.sleep(1) except Exception as e: print(f"Error sending termination message: {e}") if ws_app: ws_app.close() ws_thread.join(timeout=2.0) finally: # Final cleanup if stream and stream.is_active(): stream.stop_stream() if stream: stream.close() if audio: audio.terminate() print("Cleanup complete. Exiting.") if __name__ == "__main__": run() ``` These settings wait longer before ending turns to accommodate natural conversation pauses and ensure readable formatted text for display. You can [tweak these settings](/streaming/getting-started/transcribe-streaming-audio) to get the best results for your notetaker. ## How Do I Handle Multichannel Meeting Audio? Many meeting platforms (Zoom, Teams, Google Meet) can record each participant on separate audio channels. This dramatically improves speaker identification accuracy. ### For Pre-recorded Meetings ```python config = aai.TranscriptionConfig( multichannel=True, # Enable when each speaker is on different channel speaker_labels=False, # Disable - channels already separate speakers # Other settings... ) transcriber = aai.Transcriber() transcript = transcriber.transcribe(audio_file, config=config) # Access per-channel transcripts for channel, channel_transcript in enumerate(transcript.channels): print(f"\n=== Channel {channel} ===") print(channel_transcript.text) ``` **When to use multichannel:** - Zoom local recordings with "Record separate audio file for each participant" enabled - Professional podcast recordings with individual microphones - Conference systems with dedicated channels per participant - Phone calls with caller and callee on separate channels **Benefits:** - **Perfect speaker separation** - No diarization errors - **No speaker confusion or overlap issues** - **Faster processing time** - Diarization not needed - **Higher accuracy** - Model processes clean single-speaker audio **How to enable in meeting platforms:** - **Zoom**: Settings → Recording → Advanced → "Record a separate audio file for each participant" - **Teams**: Requires third-party recording solutions like [Recall.ai](https://www.recall.ai/) - **Google Meet**: Requires third-party recording solutions like [Recall.ai](https://www.recall.ai/) ### For Streaming Meetings For real-time multichannel audio, create separate streaming sessions per channel: ```python expandable import asyncio import websockets class ChannelTranscriber: def __init__(self, channel_id: int, speaker_name: str): self.channel_id = channel_id self.speaker_name = speaker_name self.connection_params = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "format_turns": True, } async def transcribe_channel(self, audio_stream): """Transcribe a single audio channel""" url = f"wss://streaming.assemblyai.com/v3/ws?{urlencode(self.connection_params)}" # If you're using `websockets` version 13.0 or later, use `additional_headers` parameter. For older versions (< 13.0), use `extra_headers` instead. async with websockets.connect(url, additional_headers={"Authorization": API_KEY}) as ws: # Send audio from this channel only async for audio_chunk in audio_stream: await ws.send(audio_chunk) # Receive transcripts async for message in ws: data = json.loads(message) if data.get("type") == "Turn" and data.get("end_of_turn"): print(f"{self.speaker_name}: {data['transcript']}") # Create transcriber for each channel async def transcribe_multichannel_meeting(channel_audio_streams): transcribers = [ ChannelTranscriber(0, "Alice"), ChannelTranscriber(1, "Bob"), ] # Run all channels concurrently await asyncio.gather(*[ t.transcribe_channel(stream) for t, stream in zip(transcribers, channel_audio_streams) ]) ``` See our [multichannel streaming guide](/streaming/label-speakers-and-separate-channels#multichannel-streaming-audio) for complete implementation details. ## How Should I Handle Pre-recorded Transcription in Production? Choose the right approach based on your application's needs: ### Option 1: Simple Blocking Call ```python # Simple blocking call transcript = await asyncio.to_thread(transcriber.transcribe, audio_url, config=config) ``` **Pros:** - Simple, straightforward code - Good for low volume applications - Easy to understand and debug **Cons:** - Ties up resources while waiting - Not suitable for high volume - Cannot process multiple files simultaneously **Best for:** Personal projects, prototypes, low-traffic applications ### Option 2: Webhook Callbacks (Production Recommended) ```python config = aai.TranscriptionConfig( webhook_url="https://your-app.com/webhooks/assemblyai", webhook_auth_header_name="X-Webhook-Secret", webhook_auth_header_value="your_secret_here", speaker_labels=True, summarization=True, # ... other config ) # Submit job and return immediately (non-blocking) transcript = transcriber.submit(audio_url, config=config) print(f"Job submitted: {transcript.id}") # Your app can continue processing other requests # Your webhook receives results when ready (typically 15-30% of audio duration) ``` **Webhook handler example:** ```python expandable from flask import Flask, request, jsonify app = Flask(__name__) @app.route("/webhooks/assemblyai", methods=["POST"]) def assemblyai_webhook(): # Verify webhook authenticity if request.headers.get("X-Webhook-Secret") != "your_secret_here": return jsonify({"error": "Unauthorized"}), 401 import requests as http_requests data = request.json transcript_id = data["transcript_id"] status = data["status"] if status == "completed": # Fetch the full transcript (webhook only sends transcript_id and status) transcript = http_requests.get( f"https://api.assemblyai.com/v2/transcript/{transcript_id}", headers={"authorization": "your_api_key"} ).json() process_completed_meeting(transcript) elif status == "error": log_transcription_error(transcript_id) return jsonify({"received": True}), 200 def process_completed_meeting(transcript): """Process completed meeting transcript""" utterances = transcript["utterances"] summary = transcript["summary"] # Store in database save_to_database(transcript) # Notify user send_notification(transcript["id"]) ``` **Pros:** - Non-blocking - submit and forget - Scales to high volume - Process multiple files in parallel - Automatic retry on failures - Get notified when complete **Best for:** Production apps, user-uploaded recordings, batch processing, SaaS products ### Option 3: Polling (Custom Workflows) ```python # Submit job transcript = transcriber.submit(audio_url, config=config) print(f"Submitted: {transcript.id}") # Poll for completion with progress tracking while transcript.status not in [aai.TranscriptStatus.completed, aai.TranscriptStatus.error]: await asyncio.sleep(5) transcript = transcriber.get_transcript(transcript.id) # Optional: Show progress print(f"Status: {transcript.status}...") if transcript.status == aai.TranscriptStatus.completed: process_transcript(transcript) else: print(f"Error: {transcript.error}") ``` **Pros:** - Full control over retry logic - Can show progress to users - Good for background jobs - Works without webhook infrastructure **Cons:** - Must implement your own polling loop - Ties up resources while polling - More complex than webhooks **Best for:** Background job processors, CLIs with progress bars, custom retry logic ### Comparison Table | Method | Blocking | Scalability | Complexity | Best For | | ----------- | -------- | ----------- | ---------- | ---------------------------- | | Blocking | Yes | Low | Low | Prototypes, low volume | | Webhooks | No | High | Medium | Production, high volume | | Polling | Partial | Medium | Medium | Background jobs, progress UI | ### Scaling Considerations - **HTTP rate limit:** 20,000 requests per 5-minute window, counted across submissions (POST) and polling (GET) combined - **Exceeding the limit:** returns a `403` response - **Parallel transcriptions (rate limit):** 200+ for paid accounts (queued beyond that) - **Ramp up gradually:** start at 10-50 parallel requests, double incrementally - **Avoid the rate limit:** use [webhooks](/pre-recorded-audio/webhooks) or jittered, widened polling — see [Polling without exceeding the rate limit](/pre-recorded-audio/guides/bulk-transcription-and-load-tests-at-scale#polling-without-exceeding-the-rate-limit) - **Contact Sales** before large-scale rollouts ## How Do I Identify Speakers in My Recording? Speaker diarization tells you **when** speakers change ("Speaker A", "Speaker B"), but **Speaker Identification** tells you **who** they are by name or role. ### Why Use Speaker Identification? **Instead of:** ``` Speaker A: Let's review the Q3 numbers. Speaker B: Revenue was up 15% this quarter. Speaker A: Excellent work on that launch. ``` **You get:** ``` Sarah Chen: Let's review the Q3 numbers. Michael Rodriguez: Revenue was up 15% this quarter. Sarah Chen: Excellent work on that launch. ``` ### How It Works Speaker Identification uses AssemblyAI's Speech Understanding API to map generic speaker labels to actual names or roles that you provide: ```python expandable import assemblyai as aai aai.settings.api_key = "your_api_key" # Step 1: Transcribe with speaker diarization config = aai.TranscriptionConfig( speaker_labels=True, # Must enable speaker diarization first speech_understanding={ "request": { "speaker_identification": { "speaker_type": "name", # or "role" "known_values": ["Sarah Chen", "Michael Rodriguez", "Alex Kim"] } } } ) transcriber = aai.Transcriber() transcript = transcriber.transcribe("meeting_recording.mp3", config=config) # Access results with identified speakers for utterance in transcript.utterances: print(f"{utterance.speaker}: {utterance.text}") ``` ### Identifying by Role Instead of Name For customer service, sales calls, or scenarios where you don't know names: ```python config = aai.TranscriptionConfig( speaker_labels=True, speech_understanding={ "request": { "speaker_identification": { "speaker_type": "role", "known_values": ["Agent", "Customer"] # or ["Interviewer", "Interviewee"] } } } ) ``` **Common role combinations:** - `["Agent", "Customer"]` - Customer service calls - `["Support", "Customer"]` - Technical support - `["Interviewer", "Interviewee"]` - Interviews - `["Host", "Guest"]` - Podcasts - `["Doctor", "Patient"]` - Medical consultations (with HIPAA compliance) ### How to Get Speaker Names **For platform recordings:** 1. **Zoom**: Extract participant names from Zoom API or meeting JSON 2. **Teams**: Get attendees from Microsoft Graph API 3. **Google Meet**: Use Google Calendar API to get participants **Example with Zoom:** ```python # Get participant names from Zoom meeting zoom_participants = get_zoom_meeting_participants(meeting_id) speaker_names = [p["name"] for p in zoom_participants] # Use in speaker identification config = aai.TranscriptionConfig( speaker_labels=True, speakers_expected=len(speaker_names), # Exact number of speakers to detect speech_understanding={ "request": { "speaker_identification": { "speaker_type": "name", "known_values": speaker_names } } } ) ``` ### How Speaker Identification Works **Speaker Identification Requirements:** 1. **Speaker diarization must be enabled** - Cannot identify speakers without diarization first 2. **Requires sufficient audio per speaker** - Each speaker needs enough speech for accurate matching 3. **Works best with distinct voices** - Similar voices may be confused 4. **Post-processing step** - Adds additional processing time after transcription **Accuracy depends on:** - Audio quality (clear, minimal background noise) - Voice distinctiveness (different genders, accents, tones) - Amount of speech per speaker (more = better) - Number of speakers (fewer = more accurate) ### Alternative: Add Identification Later You can add speaker identification to an existing transcript by posting to the Speech Understanding API with the `transcript_id`. This is useful when you get speaker names after the transcription completes, or when building iterative workflows where users confirm speaker identities. ```python expandable import requests # First, transcribe with speaker diarization transcript = transcriber.transcribe(audio_url, config=aai.TranscriptionConfig(speaker_labels=True)) # Later, add speaker identification using the transcript ID understanding_body = { "transcript_id": transcript.id, "speech_understanding": { "request": { "speaker_identification": { "speaker_type": "name", "known_values": ["Sarah Chen", "Michael Rodriguez"] } } } } result = requests.post( "https://llm-gateway.assemblyai.com/v1/understanding", headers={"Authorization": aai.settings.api_key}, json=understanding_body ).json() # Access identified speakers from the response for utterance in result["utterances"]: print(f"{utterance['speaker']}: {utterance['text']}") ``` This approach is useful when: - You get speaker names after the transcription completes - You want to try different name mappings - Building iterative workflows where users confirm speaker identities For complete API details, see our [Speaker Identification documentation](/speech-understanding/speaker-identification). ## How Do I Translate Between Languages in Meetings? AssemblyAI supports translation between 86 languages, enabling you to transcribe meetings in one language and translate to another. ### When to Use Translation **Common use cases:** - Transcribe Spanish meeting → Translate to English for documentation - Transcribe multilingual meeting → Translate all to common language - Create translated meeting notes for international teams - Provide translated summaries for stakeholders ### Basic Translation Translation is a Speech Understanding feature. You enable it via the `speech_understanding` parameter with `target_languages`: ```python expandable import requests import time base_url = "https://api.assemblyai.com" headers = {"authorization": "YOUR_API_KEY"} # Configure transcription with translation data = { "audio_url": "https://assembly.ai/wildfires.mp3", "speech_models": ["universal-3-5-pro", "universal-2"], "language_detection": True, "speaker_labels": True, "speech_understanding": { "request": { "translation": { "target_languages": ["es", "de"], "formal": True } } } } response = requests.post(base_url + "/v2/transcript", headers=headers, json=data) transcript_id = response.json()["id"] polling_endpoint = base_url + f"/v2/transcript/{transcript_id}" while True: transcript = requests.get(polling_endpoint, headers=headers).json() if transcript["status"] == "completed": break elif transcript["status"] == "error": raise RuntimeError(f"Transcription failed: {transcript['error']}") else: time.sleep(3) print("--- Original Transcript ---") print(transcript["text"][:200] + "...") print("\n--- Translations ---") for language_code, translated_text in transcript["translated_texts"].items(): print(f"{language_code.upper()}:") print(translated_text[:200] + "...") ``` ### Translation with Speaker Labels For meetings where you need per-utterance translations with speaker attribution: ```python data = { "audio_url": audio_url, "speech_models": ["universal-3-5-pro", "universal-2"], "speaker_labels": True, "speech_understanding": { "request": { "translation": { "target_languages": ["es"], "match_original_utterance": True, "formal": True } } } } for utterance in transcript["utterances"]: print(f"Speaker {utterance['speaker']}:") print(f" Original: {utterance['text'][:100]}...") print(f" Spanish: {utterance['translated_texts']['es'][:100]}...") ``` ### Supported Language Pairs AssemblyAI supports translation between **86 languages**, including: **Popular combinations:** - Spanish ↔ English - French ↔ English - German ↔ English - Mandarin ↔ English - Japanese ↔ English - Portuguese ↔ English - And all combinations between supported languages ### Translation Response Format The response includes `translated_texts` as a dictionary keyed by language code: ```python { "text": "Original transcript in source language", "translated_texts": { "es": "Translated transcript in Spanish", "de": "Translated transcript in German" }, "utterances": [ { "speaker": "A", "text": "Hello, how are you?", "translated_texts": { "es": "Hola, ¿cómo estás?" }, "start": 0, "end": 1500 } ] } ``` For complete language support and translation details, see our [Translation documentation](/speech-understanding/translation). ## What Workflows Can I Build for My AI Meeting Notetaker? Use these Speech Understanding and Guardrails features to transform raw transcripts into actionable insights. ### Summarization `summarization: true` **What it does:** Generates an abstractive recap of the conversation (not verbatim). **Output:** `summary` string (bullets/paragraph format). **Great for:** Meeting notes, call recaps, executive summaries. **Notes:** Condenses and rephrases; minor details may be omitted by design. **Example:** ```python config = aai.TranscriptionConfig( summarization=True, summary_type="bullets", # or "bullets_verbose", "gist", "headline", "paragraph" summary_model="informative", # or "conversational" ) ``` ### Sentiment Analysis `sentiment_analysis: true` **What it does:** Scores per-utterance sentiment (positive / neutral / negative). **Output:** Array of `{ text, sentiment, confidence, start, end }`. **Great for:** Customer satisfaction tracking, coaching, churn prediction. **Notes:** Segment-level (not global mood); sarcasm and very short utterances are harder to classify. **Example:** ```python for utterance in transcript.sentiment_analysis_results: if utterance.sentiment == "NEGATIVE": print(f"Negative sentiment detected: {utterance.text}") ``` ### Entity Detection `entity_detection: true` **What it does:** Extracts named entities (people, organizations, locations, products, etc.). **Output:** Array of `{ entity_type, text, start, end }`. **Great for:** Auto-tagging topics, tracking competitors mentioned, CRM enrichment. **Notes:** Operates on post-redaction text if PII redaction is enabled. **Example:** ```python # Extract all organizations mentioned organizations = [ entity.text for entity in transcript.entities if entity.entity_type == "organization" ] print(f"Companies mentioned: {', '.join(organizations)}") ``` ### Redact PII Text `redact_pii: true` **What it does:** Scans transcript for personally identifiable information and replaces matches per policy. **Output:** `text` with replacements; original `words` timing preserved. **Great for:** GDPR/CCPA compliance, safe sharing, SOC2 requirements. **Notes:** Runs **before** downstream features; they see the redacted text. **Recommended policies for meetings:** ```python config = aai.TranscriptionConfig( redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, # Remove names PIIRedactionPolicy.email_address, # Remove emails PIIRedactionPolicy.phone_number, # Remove phone numbers PIIRedactionPolicy.organization, # Remove company names ], redact_pii_sub=PIISubstitutionPolicy.hash, # Stable hash tokens ) ``` **Why hash substitution?** - Stable across the file (same value → same token) - Maintains sentence structure for LLM processing - Prevents reconstruction of original data ### Redact PII Audio `redact_pii_audio: true` **What it does:** Produces a second audio file where redacted portions are bleeped/silenced. **Output:** `redacted_audio_url` in the transcript response. **Great for:** External sharing, training materials, demos. **Notes:** Original audio is untouched; bleeped sections may sound choppy. ### Complete Example ```python expandable config = aai.TranscriptionConfig( # Core transcription speaker_labels=True, # Speech Understanding summarization=True, sentiment_analysis=True, entity_detection=True, # PII protection redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.email_address, PIIRedactionPolicy.phone_number, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True, ) transcript = transcriber.transcribe(audio_url, config=config) # Access all features meeting_insights = { "summary": transcript.summary, "sentiment_trend": analyze_sentiment_trend(transcript.sentiment_analysis_results), "entities": extract_entities(transcript.entities), "safe_transcript": transcript.text, # PII redacted "safe_audio": transcript.redacted_audio_url, # PII bleeped } ``` ## How Do I Improve the Accuracy of My Notetaker? **Best practices:** - Include participant names for better speaker recognition - Add company-specific jargon and acronyms - Include product names and technical terms - Keep individual terms under 50 characters - Up to 200 terms per request (Universal-2) or 1000 terms (Universal-3.5 Pro) ### Using Keyterms Prompt for Pre-recorded Transcription Keyterms prompting improves recognition accuracy for domain-specific vocabulary by up to 21%: ```python expandable # Define domain-specific vocabulary company_terms = [ "AssemblyAI", "Universal-3.5 Pro", "Speech Understanding", "diarization" ] participant_names = [ "Dylan Fox", "Sarah Chen", "Michael Rodriguez" ] technical_terms = [ "API endpoint", "WebSocket", "latency metrics", "TTFT" ] # Configure with keyterms prompt config = aai.TranscriptionConfig( keyterms_prompt=company_terms + participant_names + technical_terms, speaker_labels=True, # ... other settings ) ``` ### Using Keyterms Prompt for Streaming ```python expandable # Streaming with contextual keyterms keyterms = [ # Participant names "Alice Johnson", "Bob Smith", # Meeting-specific vocabulary "Q4 objectives", "revenue targets", "customer acquisition", # Technical terms "API integration", "cloud migration" ] CONNECTION_PARAMS = { "sample_rate": 16000, "speech_model": "universal-3-5-pro", "format_turns": True, "keyterms_prompt": keyterms, } ``` ## How Do I Process the Response from the API? ### Processing Pre-recorded Responses ```python expandable def process_transcript(transcript): """ Extract and process all relevant data from pre-recorded transcript """ # Basic transcript data meeting_data = { "id": transcript.id, "duration": transcript.audio_duration, "confidence": transcript.confidence, "full_text": transcript.text } # Process speaker utterances speakers = {} for utterance in transcript.utterances: speaker = utterance.speaker if speaker not in speakers: speakers[speaker] = { "utterances": [], "total_speaking_time": 0, "word_count": 0 } speakers[speaker]["utterances"].append({ "text": utterance.text, "start": utterance.start, "end": utterance.end, "confidence": utterance.confidence }) # Calculate speaking time speakers[speaker]["total_speaking_time"] += (utterance.end - utterance.start) / 1000 speakers[speaker]["word_count"] += len(utterance.text.split()) meeting_data["speakers"] = speakers # Extract summary if transcript.summary: meeting_data["summary"] = transcript.summary # Calculate meeting statistics total_duration = transcript.audio_duration # Already in seconds meeting_data["statistics"] = { "total_speakers": len(speakers), "total_words": sum(s["word_count"] for s in speakers.values()), "average_confidence": transcript.confidence, "speaking_distribution": { speaker: { "percentage": (data["total_speaking_time"] / total_duration) * 100, "minutes": data["total_speaking_time"] / 60 } for speaker, data in speakers.items() } } return meeting_data # Example usage result = process_transcript(transcript) print(f"Meeting had {result['statistics']['total_speakers']} speakers") print(f"Speaker A spoke for {result['statistics']['speaking_distribution']['A']['minutes']:.1f} minutes") ``` ### Processing Streaming Responses ```python expandable class StreamingResponseProcessor: def __init__(self): self.partial_buffer = "" self.final_transcripts = [] self.turn_metadata = [] def process_message(self, message: dict): """ Process real-time streaming messages """ msg_type = message.get("type") if msg_type == "Begin": return { "event": "session_started", "session_id": message.get("id"), "expires_at": message.get("expires_at") } elif msg_type == "Turn": return self.process_turn(message) elif msg_type == "Termination": return { "event": "session_ended", "audio_duration": message.get("audio_duration_seconds"), "session_duration": message.get("session_duration_seconds") } def process_turn(self, data: dict): """Process turn messages""" is_final = data.get("end_of_turn") transcript = data.get("transcript", "") turn_order = data.get("turn_order") response = { "turn_order": turn_order, "is_final": is_final, "confidence": data.get("end_of_turn_confidence", 0) } # Handle partials (for live display) if not is_final and transcript: self.partial_buffer = transcript response["event"] = "partial" response["text"] = transcript # Handle finals (for storage) elif is_final: final_transcript = { "turn_order": turn_order, "text": transcript, "confidence": data.get("end_of_turn_confidence"), "timestamp": datetime.now().isoformat() } self.final_transcripts.append(final_transcript) response["event"] = "final" response["text"] = transcript # Clear partial buffer self.partial_buffer = "" return response def get_full_transcript(self): """ Combine all final transcripts into complete meeting transcript """ return { "full_text": " ".join(t["text"] for t in self.final_transcripts), "transcripts": self.final_transcripts, "total_turns": len(self.final_transcripts) } # Example usage processor = StreamingResponseProcessor() # If you're using `websockets` version 13.0 or later, use `additional_headers` parameter. For older versions (< 13.0), use `extra_headers` instead. async with websockets.connect(API_ENDPOINT, additional_headers=headers) as ws: async for message in ws: data = json.loads(message) result = processor.process_message(data) if result["event"] == "partial": # Update UI with live transcript update_live_caption(result["text"]) elif result["event"] == "final": # Save final transcript save_transcript_segment(result) # Get complete transcript when done full_transcript = processor.get_full_transcript() ``` ## Additional Resources - [Universal Pre-recorded Documentation](/pre-recorded-audio) - [Universal-3.5 Pro Streaming Documentation](/streaming) - [Speaker Diarization Guide](/pre-recorded-audio/label-speakers) - [Speaker Identification Guide](/speech-understanding/speaker-identification) - [Translation Guide](/speech-understanding/translation) - [Getting Started Guide](/pre-recorded-audio/getting-started/transcribe-an-audio-file) - [API Playground](https://www.assemblyai.com/playground/streaming) - [Changelog](https://www.assemblyai.com/changelog) - [Support](https://www.assemblyai.com/contact/support) --- # Build a medical scribe URL: https://www.assemblyai.com/docs/medical-scribe-best-practices Source: docs/medical-scribe-best-practices.mdx Navigation: Pre-recorded STT > Guides > Tutorials Description: Build AI-powered medical scribes that transcribe patient encounters and generate clinical documentation. AssemblyAI provides everything you need to build a medical scribe, from high-accuracy transcription with medical terminology support to HIPAA-compliant PII redaction and structured clinical note generation through the LLM Gateway. Choose the guide that matches your clinical workflow: Transcribe recorded patient encounters with Universal-3.5 Pro. Includes Medical Mode, speaker diarization, entity detection, PII redaction, and SOAP note generation. Stream audio from a microphone during live encounters with Universal-3.5 Pro Streaming. Includes LLM Gateway post-processing, medical keyterms, and automatic SOAP note generation. ## Which approach should I use? | Scenario | Recommended approach | | --- | --- | | Post-visit documentation with highest accuracy | [Post-visit medical scribe](/medical-scribe-best-practices/medical-scribe-post-visit) | | Live encounter transcription during patient visits | [Real-time medical scribe](/medical-scribe-best-practices/medical-scribe-real-time) | | Telemedicine or emergency department visits | [Real-time medical scribe](/medical-scribe-best-practices/medical-scribe-real-time) | | Specialist consultations with complex terminology | [Post-visit medical scribe](/medical-scribe-best-practices/medical-scribe-post-visit) | | Hybrid: real-time notes with post-visit verification | Use both guides together | --- # Best Practices for building Contact Center Applications URL: https://www.assemblyai.com/docs/contact-center-best-practices Source: docs/contact-center-best-practices.mdx Navigation: Pre-recorded STT > Guides > Tutorials Description: Complete guide for building contact center applications with AssemblyAI ## Introduction Building a contact center application requires careful consideration of accuracy, speaker separation, compliance, and scalability. This guide addresses common questions and provides practical solutions for both post-call analytics and real-time agent assist scenarios. ## Why AssemblyAI for contact centers? AssemblyAI stands out as the premier choice for contact center applications with several key advantages: ### Industry-leading accuracy on telephony audio - **Universal-3.5 Pro model** delivers best-in-class accuracy on 8kHz telephony audio - **2.9% speaker diarization error rate** for precise agent vs. customer attribution - **Multichannel support** for stereo call recordings where agent and customer are on separate channels - **Keyterms prompt** allows providing call context to improve accuracy of company names, products, and compliance phrases ### Streaming with Universal-3.5 Pro For real-time agent assist, AssemblyAI's Universal-3.5 Pro Streaming model (`universal-3-5-pro`) offers: - **Low latency** enables live transcription during calls - **Format turns** feature provides structured, readable output - **Dynamic prompting** via `UpdateConfiguration` to update context mid-call - **Dual-channel streaming** for separate agent and customer audio streams ### End-to-end voice AI platform Unlike fragmented solutions, AssemblyAI provides a unified API for: - Transcription with speaker diarization (agent vs. customer) - Multichannel audio support for stereo call recordings - PII redaction on both text and audio for HIPAA and PCI compliance - Post-processing workflows with custom prompting - from call summaries to QA scoring - Streaming and pre-recorded transcription in a single platform - Compliance and security built for enterprise workloads (BAA, SOC2, ISO) ## When should I use pre-recorded vs streaming for contact centers? Understanding when to use pre-recorded versus streaming is critical for contact center workflows. ### Pre-recorded Speech-to-text **Post-call analytics** - Call already happened, you have the full recording - **Highest accuracy needed** - Pre-recorded models have the highest accuracy - **Speaker diarization is critical** - Pre-recorded has 2.9% speaker error rate - **Multichannel recordings** - Most contact center recordings are stereo with agent and customer on separate channels - **Compliance workflows** - Full PII redaction with audio de-identification - **Post-call analytics** - Summarization, sentiment analysis, entity detection, QA scoring - **Batch processing** - Processing large volumes of call recordings **Best for:** QA scoring, compliance monitoring, coaching insights, post-call CRM updates, searchable call archives ### Streaming Speech-to-text **Live calls** - Transcribing as the call happens You should use streaming when you need to display a live transcript to agents during calls. With Universal-3.5 Pro Streaming, accuracy is closer to pre-recorded, but pre-recorded will always be the most accurate option. - **Agent assist** - Live transcription visible to agents during calls - **Real-time coaching** - Prompt agents with suggested responses or compliance reminders - **Live compliance monitoring** - Detect compliance violations in real-time - **No recording available** - Processing live audio only **Best for:** Agent assist, real-time coaching, live compliance monitoring, live call transcription ### Hybrid approach (recommended) Many contact center platforms use **both**: 1. **Streaming during the call** - Provide live transcription for agent assist and real-time coaching 2. **Pre-recorded after the call** - Generate high-quality transcript with speaker labels, summary, and analytics **Example workflow:** - Call begins → Start streaming for live agent assist - Call ends → Upload recording to pre-recorded API for final transcript with speaker names - Generate call summary, QA score, and compliance report from pre-recorded transcript - Push results to CRM (e.g., Salesforce) ## What languages and features for a contact center application? ### Pre-recorded calls (Universal-3.5 Pro) For post-call analytics, AssemblyAI supports: **Languages**: - 99 languages supported - Automatic Language Detection to route to the most spoken language - Code Switching to preserve changes in speech between languages **Core Features**: - Speaker diarization (agent-customer separation) - Multichannel audio support - when agent and customer are on separate audio channels, enables perfect speaker separation without diarization - Automatic formatting, punctuation, and capitalization - Keyterms prompting for boosting domain-specific terms (up to 1000 terms for Universal-3.5 Pro) - Natural language prompting (Universal-3.5 Pro) - up to 1,500 words to guide transcription behavior - Speaker options with configurable min/max expected speakers for call transfers **Speech Understanding**: - Summarization for call recaps - Sentiment analysis for customer satisfaction tracking - Entity detection for extracting names, account numbers, and products - Speaker identification to map generic labels to agent and customer names - Translation between 86 languages **Guardrails**: - PII redaction on text and audio for HIPAA and PCI compliance ### Streaming (Universal-3.5 Pro Streaming) For live call transcription, use **Universal-3.5 Pro Streaming** (`universal-3-5-pro`) for the highest streaming accuracy: **Core Features**: - Speaker diarization for identifying agent vs. customer - Partial and final transcripts for responsive UI - Format turns for structured, readable output - Keyterms prompt for company names, products, and compliance phrases - Dual-channel streaming for separate agent and customer audio For more details, see the [Universal-3.5 Pro Streaming documentation](/streaming/getting-started/transcribe-streaming-audio). ## How can I get started building a post-call analytics pipeline? Here's a complete example implementing pre-recorded transcription for contact center call analysis: ```python expandable import assemblyai as aai import asyncio from typing import Dict, List from assemblyai.types import ( SpeakerOptions, PIIRedactionPolicy, PIISubstitutionPolicy, ) # Configure API key aai.settings.api_key = "your_api_key_here" async def transcribe_call(audio_source: str, agent_name: str = None) -> Dict: """ Transcribe a contact center call recording with full analytics Args: audio_source: Either a local file path or publicly accessible URL agent_name: Optional agent name for speaker identification """ # Configure comprehensive call analysis config = aai.TranscriptionConfig( # Model selection speech_models=["universal-3-5-pro", "universal-2"], # Speaker diarization speaker_labels=True, speaker_options=SpeakerOptions( min_speakers_expected=2, # Agent and customer max_speakers_expected=5 # Allow for call transfers - safe to keep high ), multichannel=False, # Set to True if audio has separate channel per speaker # Language detection language_detection=True, # Boost accuracy of contact center vocabulary keyterms_prompt=[ # Company-specific terms "Acme Corp", "Premium Support Plan", # Compliance phrases "recorded line", "calls are monitored and recorded", # Common contact center terms "account number", "case number", "ticket number", "escalation", "supervisor", "hold time", ], # Post-call analytics summarization=True, sentiment_analysis=True, entity_detection=True, # PII protection for compliance redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.phone_number, PIIRedactionPolicy.email_address, PIIRedactionPolicy.account_number, PIIRedactionPolicy.us_social_security_number, PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.credit_card_cvv, PIIRedactionPolicy.credit_card_expiration, PIIRedactionPolicy.date_of_birth, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True, ) # Add speaker identification if agent name is known if agent_name: config.speech_understanding = { "request": { "speaker_identification": { "speaker_type": "role", "speakers": [ {"role": "Agent", "name": agent_name}, {"role": "Customer"} ] } } } # Create transcriber transcriber = aai.Transcriber() try: # Submit transcription job transcript = await asyncio.to_thread( transcriber.transcribe, audio_source, config=config ) # Check status if transcript.status == aai.TranscriptStatus.error: raise Exception(f"Transcription failed: {transcript.error}") # Process speaker-labeled utterances for utterance in transcript.utterances: start_time = utterance.start / 1000 # Convert ms to seconds end_time = utterance.end / 1000 print(f"[{start_time:.1f}s - {end_time:.1f}s] {utterance.speaker}:") print(f" {utterance.text}\n") return { "transcript": transcript, "utterances": transcript.utterances, "summary": transcript.summary, "sentiment": transcript.sentiment_analysis_results, "entities": transcript.entities, "redacted_audio_url": transcript.redacted_audio_url, } except Exception as e: print(f"Error during transcription: {e}") raise async def main(): audio_source = "https://your-storage.com/calls/call_recording.mp3" result = await transcribe_call(audio_source, agent_name="Sarah Johnson") print(f"\nCall duration: {result['transcript'].audio_duration} seconds") print(f"Summary: {result['summary']}") if __name__ == "__main__": asyncio.run(main()) ``` ## How Do I Handle Multichannel Contact Center Audio? Most contact center recordings are stereo with the agent on one channel and the customer on the other. Multichannel transcription gives you perfect speaker separation without diarization. ### Pre-recorded Multichannel ```python expandable config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], multichannel=True, # Enable when agent and customer are on separate channels speaker_labels=False, # Disable - channels already separate speakers # Still enable analytics summarization=True, sentiment_analysis=True, entity_detection=True, # PII redaction redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.us_social_security_number, PIIRedactionPolicy.account_number, ], redact_pii_sub=PIISubstitutionPolicy.hash, ) transcriber = aai.Transcriber() transcript = transcriber.transcribe(audio_file, config=config) # Channel 1 = Agent, Channel 2 = Customer (typical layout) for utterance in transcript.utterances: role = "Agent" if utterance.channel == "1" else "Customer" print(f"{role}: {utterance.text}") ``` **When to use multichannel:** - Call recordings from PBX systems with separate agent/customer channels - Recordings from platforms like Genesys, Twilio, Five9, NICE, or Talkdesk - Any stereo recording where each channel represents a different speaker **Benefits:** - **Perfect speaker separation** - No diarization errors - **No speaker confusion or overlap issues** - **Higher accuracy** - Model processes clean single-speaker audio per channel ### Streaming Multichannel For real-time dual-channel transcription, create separate streaming sessions per channel: ```python expandable import asyncio import websockets import json from urllib.parse import urlencode API_KEY = "your_api_key" class ChannelTranscriber: def __init__(self, channel_id: int, role: str): self.channel_id = channel_id self.role = role self.connection_params = { "sample_rate": 8000, # Telephony standard "speech_model": "universal-3-5-pro", "format_turns": True, "encoding": "pcm_mulaw", # Common telephony encoding } async def transcribe_channel(self, audio_stream): url = f"wss://streaming.assemblyai.com/v3/ws?{urlencode(self.connection_params, doseq=True)}" # If using websockets >= 13.0, use additional_headers. For < 13.0, use extra_headers. async with websockets.connect(url, additional_headers={"Authorization": API_KEY}) as ws: # Send and receive must run concurrently for real-time streaming async def send_audio(): async for audio_chunk in audio_stream: await ws.send(audio_chunk) async def receive_transcripts(): async for message in ws: data = json.loads(message) if data.get("type") == "Turn" and data.get("end_of_turn"): print(f"{self.role}: {data['transcript']}") await asyncio.gather(send_audio(), receive_transcripts()) # Create transcriber for each channel async def transcribe_live_call(agent_audio_stream, customer_audio_stream): agent = ChannelTranscriber(0, "Agent") customer = ChannelTranscriber(1, "Customer") await asyncio.gather( agent.transcribe_channel(agent_audio_stream), customer.transcribe_channel(customer_audio_stream), ) ``` See our [multichannel streaming guide](/streaming/label-speakers-and-separate-channels#multichannel-streaming-audio) for complete implementation details. ## How Can I Build a Real-Time Agent Assist? Here's a complete example for real-time streaming transcription optimized for contact center agent assist: ```python expandable # pip install pyaudio websocket-client import pyaudio import websocket import json import threading import time from urllib.parse import urlencode from datetime import datetime # --- Configuration --- YOUR_API_KEY = "your_api_key" # Contact center keyterms KEYTERMS = [ # Company and product terms "Acme Corp", "Premium Support Plan", "Enterprise License", # Compliance phrases "recorded line", "calls are monitored", # Common contact center vocabulary "account number", "case number", "escalation", "supervisor", ] # CONTACT CENTER CONFIGURATION CONNECTION_PARAMS = { "sample_rate": 8000, # Telephony standard (8kHz) "speech_model": "universal-3-5-pro", # Universal-3.5 Pro Streaming for highest accuracy "format_turns": True, # Contact center turn detection # universal-3-5-pro defaults: min_turn_silence=100ms, max_turn_silence=1000ms "min_turn_silence": 400, # Longer than default for natural call pauses "max_turn_silence": 1500, # Longer for customers explaining issues # Keyterms for accuracy "keyterms_prompt": KEYTERMS, } API_ENDPOINT_BASE_URL = "wss://streaming.assemblyai.com/v3/ws" API_ENDPOINT = f"{API_ENDPOINT_BASE_URL}?{urlencode(CONNECTION_PARAMS, doseq=True)}" # Audio Configuration FRAMES_PER_BUFFER = 400 # 50ms of audio at 8kHz SAMPLE_RATE = CONNECTION_PARAMS["sample_rate"] CHANNELS = 1 FORMAT = pyaudio.paInt16 # Global variables audio = None stream = None ws_app = None audio_thread = None stop_event = threading.Event() transcript_buffer = [] def on_open(ws): print("=" * 80) print(f"[{datetime.now().strftime('%H:%M:%S')}] Agent assist transcription started") print(f"Connected to: {API_ENDPOINT_BASE_URL}") print(f"Keyterms configured: {', '.join(KEYTERMS[:5])}...") print("=" * 80) def stream_audio(): global stream while not stop_event.is_set(): try: audio_data = stream.read(FRAMES_PER_BUFFER, exception_on_overflow=False) ws.send(audio_data, websocket.ABNF.OPCODE_BINARY) except Exception as e: if not stop_event.is_set(): print(f"Error streaming audio: {e}") break global audio_thread audio_thread = threading.Thread(target=stream_audio) audio_thread.daemon = True audio_thread.start() def on_message(ws, message): try: data = json.loads(message) msg_type = data.get("type") if msg_type == "Begin": session_id = data.get("id", "N/A") print(f"[SESSION] Started - ID: {session_id}\n") elif msg_type == "Turn": end_of_turn = data.get("end_of_turn", False) transcript = data.get("transcript", "") turn_order = data.get("turn_order", 0) # Show partials for responsive agent UI if not end_of_turn and transcript: print(f"\r[LIVE] {transcript}", end="", flush=True) # Use formatted finals for agent display if end_of_turn and transcript: timestamp = datetime.now().strftime('%H:%M:%S') print(f"\n[{timestamp}] {transcript}") # Detect compliance keywords transcript_lower = transcript.lower() if any(term in transcript_lower for term in ["cancel", "refund", "complaint", "supervisor"]): print(" ** ESCALATION KEYWORD DETECTED **") transcript_buffer.append({ "timestamp": timestamp, "text": transcript, "turn_order": turn_order, "type": "final" }) print() elif msg_type == "Termination": audio_duration = data.get("audio_duration_seconds", 0) print(f"\n[SESSION] Terminated - Duration: {audio_duration}s") elif msg_type == "Error": error_msg = data.get("error", "Unknown error") print(f"\n[ERROR] {error_msg}") except json.JSONDecodeError as e: print(f"Error decoding message: {e}") except Exception as e: print(f"Error handling message: {e}") def on_error(ws, error): print(f"\n[WEBSOCKET ERROR] {error}") stop_event.set() def on_close(ws, close_status_code, close_msg): print(f"\n[WEBSOCKET] Disconnected - Status: {close_status_code}, Message: {close_msg}") global stream, audio stop_event.set() if stream: if stream.is_active(): stream.stop_stream() stream.close() stream = None if audio: audio.terminate() audio = None if audio_thread and audio_thread.is_alive(): audio_thread.join(timeout=1.0) def run(): global audio, stream, ws_app audio = pyaudio.PyAudio() try: stream = audio.open( input=True, frames_per_buffer=FRAMES_PER_BUFFER, channels=CHANNELS, format=FORMAT, rate=SAMPLE_RATE, ) except Exception as e: print(f"Error opening audio stream: {e}") if audio: audio.terminate() return ws_app = websocket.WebSocketApp( API_ENDPOINT, header={"Authorization": YOUR_API_KEY}, on_open=on_open, on_message=on_message, on_error=on_error, on_close=on_close, ) ws_thread = threading.Thread(target=ws_app.run_forever) ws_thread.daemon = True ws_thread.start() try: while ws_thread.is_alive(): time.sleep(0.1) except KeyboardInterrupt: print("\n\nCtrl+C received. Stopping transcription...") stop_event.set() if ws_app and ws_app.sock and ws_app.sock.connected: try: terminate_message = {"type": "Terminate"} ws_app.send(json.dumps(terminate_message)) time.sleep(1) except Exception as e: print(f"Error sending termination message: {e}") if ws_app: ws_app.close() ws_thread.join(timeout=2.0) finally: if stream and stream.is_active(): stream.stop_stream() if stream: stream.close() if audio: audio.terminate() print("Cleanup complete. Exiting.") if __name__ == "__main__": run() ``` ## How Should I Handle Pre-recorded Transcription in Production? ### Webhook Callbacks (Recommended) For high-volume contact center workloads, use webhooks instead of polling: ```python expandable config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], webhook_url="https://your-app.com/webhooks/assemblyai", webhook_auth_header_name="X-Webhook-Secret", webhook_auth_header_value="your_secret_here", speaker_labels=True, multichannel=True, summarization=True, sentiment_analysis=True, entity_detection=True, redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.us_social_security_number, PIIRedactionPolicy.account_number, ], redact_pii_sub=PIISubstitutionPolicy.hash, ) # Submit job and return immediately (non-blocking) transcript = transcriber.submit(audio_url, config=config) print(f"Job submitted: {transcript.id}") # Your app continues processing other calls ``` **Webhook handler example:** ```python expandable from flask import Flask, request, jsonify app = Flask(__name__) @app.route("/webhooks/assemblyai", methods=["POST"]) def assemblyai_webhook(): if request.headers.get("X-Webhook-Secret") != "your_secret_here": return jsonify({"error": "Unauthorized"}), 401 import requests as http_requests data = request.json transcript_id = data["transcript_id"] status = data["status"] if status == "completed": # Fetch the full transcript (webhook only sends transcript_id and status) transcript = http_requests.get( f"https://api.assemblyai.com/v2/transcript/{transcript_id}", headers={"authorization": "your_api_key"} ).json() process_completed_call(transcript) elif status == "error": log_transcription_error(transcript_id) return jsonify({"received": True}), 200 def process_completed_call(transcript): """Process completed call transcript and push to CRM""" utterances = transcript["utterances"] summary = transcript["summary"] # Store in database save_to_database(transcript) # Push summary to CRM push_to_crm(transcript["id"], summary) # Run QA scoring qa_score = score_call_quality(utterances) save_qa_score(transcript["id"], qa_score) ``` ### Scaling Considerations - **HTTP rate limit:** 20,000 requests per 5-minute window, counted across submissions (POST) and polling (GET) combined - **Exceeding the limit:** returns a `403` response - **Parallel transcriptions (rate limit):** 200+ for paid accounts (queued beyond that) - **Ramp up gradually:** start at 10-50 parallel requests, double incrementally - **Avoid the rate limit:** use [webhooks](/pre-recorded-audio/webhooks) or jittered, widened polling — see [Polling without exceeding the rate limit](/pre-recorded-audio/guides/bulk-transcription-and-load-tests-at-scale#polling-without-exceeding-the-rate-limit) - **Contact Sales** before large-scale rollouts ## How Do I Handle PII and Compliance? PII redaction is critical for contact center compliance (HIPAA, PCI-DSS, GDPR, CCPA). ### Recommended PII Configuration ```python expandable config = aai.TranscriptionConfig( redact_pii=True, redact_pii_policies=[ # Customer identity PIIRedactionPolicy.person_name, PIIRedactionPolicy.date_of_birth, PIIRedactionPolicy.us_social_security_number, # Contact information PIIRedactionPolicy.phone_number, PIIRedactionPolicy.email_address, PIIRedactionPolicy.location, # Financial information (PCI-DSS) PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.credit_card_cvv, PIIRedactionPolicy.credit_card_expiration, PIIRedactionPolicy.account_number, PIIRedactionPolicy.banking_information, ], redact_pii_sub=PIISubstitutionPolicy.hash, # Stable hash tokens redact_pii_audio=True, # Create de-identified audio file ) ``` **Why hash substitution?** - Stable across the file (same value = same token) - Maintains sentence structure for downstream LLM processing - Prevents reconstruction of original data ### HIPAA Compliance - AssemblyAI provides a **Business Associate Agreement (BAA)** at no cost - Paid customers can review and sign our standard online BAA self-serve from the [**Data Controls** page](https://www.assemblyai.com/dashboard/settings/data-controls) in the dashboard (Owner or Admin role required, upgraded account required) - Use PII redaction with audio de-identification for full compliance ## How Do I Improve the Accuracy of My Contact Center Transcription? ### Prompting Best Practices The most impactful lever for contact center accuracy is **prompting**. Use a structured prompt with a `Context:` field: ```python expandable config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], # Natural language prompt for transcription guidance prompt=( "Transcribe this audio with perfect punctuation and formatting. " "Preserve linguistic speech patterns including disfluencies, filler words, " "hesitations, repetitions, stutters, false starts, and colloquialisms. " "Transcribe in the original language mix (code-switching), preserving the " "words in the language they are spoken. Output plain transcript text only. " "Use a new line when the voice changes; each line contains only one " "person's words.\n\n" "Context: Acme Corp customer service call, recorded line, " "Agent: Sarah Johnson, calls are monitored and recorded" ), # Keyterms for proper nouns and domain vocabulary keyterms_prompt=[ "Acme Corp", "Sarah Johnson", "Premium Support Plan", "Enterprise License", "recorded line", "calls are monitored and recorded", ], speaker_labels=True, ) ``` **Tips for effective prompting:** - Use **positive instructions** ("transcribe verbatim") not negative ("do NOT summarize") - **Start with fewer instructions, add one at a time** — every added instruction risks conflicting with another. Treat the older "3–6 instructions" guidance as an upper bound, not a target. - **Layer instructions one by one** and test each against your call recordings to measure impact - Dynamize the `Context:` line per call with known info: company name, agent name, compliance phrases - Use **keyterms** for proper nouns and domain vocabulary (company names, product names, agent names) ### Using Keyterms for Pre-recorded Transcription ```python expandable # Build keyterms dynamically per call call_keyterms = [ # Company terms (static) "Acme Corp", "Premium Support Plan", # Agent name (from routing system) agent_name, # Customer name (from CRM lookup) customer_name, # Account-specific terms "account ending in 4532", ] config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], keyterms_prompt=call_keyterms, speaker_labels=True, ) ``` ### Using Keyterms for Streaming ```python # Streaming with contact center context keyterms = [ "Acme Corp", "Premium Support Plan", "Sarah Johnson", "recorded line", ] CONNECTION_PARAMS = { "sample_rate": 8000, "speech_model": "universal-3-5-pro", "format_turns": True, "encoding": "pcm_mulaw", "keyterms_prompt": keyterms, } ``` ## What Workflows Can I Build for My Contact Center Application? Use these features to transform raw call transcripts into actionable insights. ### Summarization `summarization: true` **What it does:** Generates an abstractive recap of the call. **Output:** `summary` string (bullets/paragraph format). **Great for:** Post-call CRM updates, call recaps, supervisor review. ```python config = aai.TranscriptionConfig( summarization=True, summary_type="bullets", # or "bullets_verbose", "gist", "headline", "paragraph" summary_model="informative", # or "conversational" ) ``` ### Sentiment Analysis `sentiment_analysis: true` **What it does:** Scores per-utterance sentiment (positive / neutral / negative). **Output:** Array of `{ text, sentiment, confidence, start, end }`. **Great for:** Customer satisfaction tracking, escalation detection, QA scoring. ```python # Analyze customer sentiment across a call negative_count = 0 for result in transcript.sentiment_analysis_results: if result.sentiment == "NEGATIVE": negative_count += 1 print(f"Negative at {result.start / 1000:.1f}s: {result.text}") # Flag calls with high negative sentiment if negative_count > 3: flag_for_supervisor_review(transcript.id) ``` ### Entity Detection `entity_detection: true` **What it does:** Extracts named entities (people, organizations, locations, products, etc.). **Output:** Array of `{ entity_type, text, start, end }`. **Great for:** CRM enrichment, auto-tagging topics, competitor tracking. ```python # Extract key entities from a call organizations = [e.text for e in transcript.entities if e.entity_type == "organization"] print(f"Companies mentioned: {', '.join(organizations)}") ``` ### Speaker Identification Map generic speaker labels to agent and customer names: ```python config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], speaker_labels=True, speech_understanding={ "request": { "speaker_identification": { "speaker_type": "role", "speakers": [ {"role": "Agent", "name": "Sarah Johnson", "description": "Customer service representative"}, {"role": "Customer"} ] } } } ) ``` ### Translation Translate call transcripts for international teams: ```python config = aai.TranscriptionConfig( speech_models=["universal-3-5-pro", "universal-2"], language_detection=True, speaker_labels=True, speech_understanding={ "request": { "translation": { "target_languages": ["en"], # Translate to English "match_original_utterance": True, # Per-utterance translations "formal": True } } } ) ``` ### Redact PII Text and Audio ```python config = aai.TranscriptionConfig( redact_pii=True, redact_pii_policies=[ PIIRedactionPolicy.person_name, PIIRedactionPolicy.credit_card_number, PIIRedactionPolicy.us_social_security_number, PIIRedactionPolicy.account_number, ], redact_pii_sub=PIISubstitutionPolicy.hash, redact_pii_audio=True, # Generate de-identified audio ) # After transcription print(transcript.text) # PII redacted in text print(transcript.redacted_audio_url) # PII bleeped in audio ``` ## How Do I Process the Response from the API? ### Processing Pre-recorded Responses ```python expandable def process_call_transcript(transcript): """ Extract and process all relevant data from a pre-recorded call transcript """ call_data = { "id": transcript.id, "duration": transcript.audio_duration, # Already in seconds "confidence": transcript.confidence, "full_text": transcript.text, } # Process speaker utterances speakers = {} for utterance in transcript.utterances: speaker = utterance.speaker if speaker not in speakers: speakers[speaker] = { "utterances": [], "total_speaking_time": 0, "word_count": 0 } speakers[speaker]["utterances"].append({ "text": utterance.text, "start": utterance.start, "end": utterance.end, }) speakers[speaker]["total_speaking_time"] += (utterance.end - utterance.start) / 1000 speakers[speaker]["word_count"] += len(utterance.text.split()) call_data["speakers"] = speakers # Extract summary if transcript.summary: call_data["summary"] = transcript.summary # Analyze sentiment if transcript.sentiment_analysis_results: sentiments = [r.sentiment for r in transcript.sentiment_analysis_results] call_data["sentiment_breakdown"] = { "positive": sentiments.count("POSITIVE"), "neutral": sentiments.count("NEUTRAL"), "negative": sentiments.count("NEGATIVE"), } # Calculate statistics total_duration = transcript.audio_duration call_data["statistics"] = { "total_speakers": len(speakers), "total_words": sum(s["word_count"] for s in speakers.values()), "speaking_distribution": { speaker: { "percentage": (data["total_speaking_time"] / total_duration) * 100, "minutes": data["total_speaking_time"] / 60, } for speaker, data in speakers.items() }, } return call_data result = process_call_transcript(transcript) print(f"Call had {result['statistics']['total_speakers']} speakers") print(f"Sentiment: {result.get('sentiment_breakdown', {})}") ``` ## Additional Resources - [Universal-3.5 Pro Pre-recorded Documentation](/pre-recorded-audio) - [Universal-3.5 Pro Streaming Documentation](/streaming/getting-started/transcribe-streaming-audio) - [Speaker Diarization Guide](/pre-recorded-audio/label-speakers) - [Multichannel Streaming Guide](/streaming/label-speakers-and-separate-channels) - [Speaker Identification Guide](/speech-understanding/speaker-identification) - [Translation Guide](/speech-understanding/translation) - [PII Redaction Guide](/guardrails/redact-pii-from-transcripts) - [Getting Started Guide](/pre-recorded-audio/getting-started/transcribe-an-audio-file) - [API Playground](https://www.assemblyai.com/playground/streaming) - [Changelog](https://www.assemblyai.com/changelog) - [Support](https://www.assemblyai.com/contact/support) --- # Get YouTube Video Transcripts with yt-dlp URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/transcribe_youtube_videos Source: docs/pre-recorded-audio/guides/transcribe_youtube_videos.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Basic transcription workflows Description: Get YouTube Video Transcripts with yt-dlp documentation. In this guide, we'll show you how to transcribe YouTube videos. For this, we use the [yt-dlp](https://github.com/yt-dlp/yt-dlp) library to download YouTube videos and then transcribe it with the AssemblyAI API. `yt-dlp` is a [youtube-dl](https://github.com/ytdl-org/youtube-dl) fork with additional features and fixes. It is better maintained and preferred over `youtube-dl` nowadays. In this guide we'll show 2 different approaches: - Option 1: Download video via CLI - Option 2: Download video via code Let's get started! ## Quickstart ```python expandable import assemblyai as aai import yt_dlp def transcribe_youtube_video(video_url: str, api_key: str) -> str: """ Transcribe a YouTube video given its URL. Args: video_url: The YouTube video URL to transcribe api_key: AssemblyAI API key Returns: The transcript text """ # Configure yt-dlp options for audio extraction ydl_opts = { 'format': 'm4a/bestaudio/best', 'outtmpl': '%(id)s.%(ext)s', 'postprocessors': [{ 'key': 'FFmpegExtractAudio', 'preferredcodec': 'm4a', }] } # Download and extract audio with yt_dlp.YoutubeDL(ydl_opts) as ydl: ydl.download([video_url]) # Get video ID from info dict info = ydl.extract_info(video_url, download=False) video_id = info['id'] # Configure AssemblyAI aai.settings.api_key = api_key # Transcribe the downloaded audio file config = aai.TranscriptionConfig() transcriber = aai.Transcriber() transcript = transcriber.transcribe(f"{video_id}.m4a", config) return transcript.text transcript_text = transcribe_youtube_video("https://www.youtube.com/watch?v=wtolixa9XTg", "YOUR-API-KEY") print(transcript_text) ``` # Step-by-step guide ## Install Dependencies Install [yt-dlp](https://github.com/yt-dlp/yt-dlp) and the [AssemblyAI Python SDK](https://github.com/AssemblyAI/assemblyai-python-sdk) via pip. ```bash pip install -U yt-dlp ``` ```bash pip install assemblyai ``` ## Option 1: Download video via CLI In this approach we download the YouTube video via the command line and then transcribe it via the AssemblyAI API. We use the following video here: - https://www.youtube.com/watch?v=wtolixa9XTg To download it, use the `yt-dlp` command with the following options: - `-f m4a/bestaudio`: The format should be the best audio version in m4a format. - `-o "%(id)s.%(ext)s"`: The output name should be the id followed by the extension. In this example, the video gets saved to "wtolixa9XTg.m4a". - `wtolixa9XTg`: the id of the video. ```bash yt-dlp -f m4a/bestaudio -o "%(id)s.%(ext)s" https://www.youtube.com/watch?v=wtolixa9XTg ``` Next, set up the AssemblyAI SDK and trancribe the file. Replace `YOUR_API_KEY` with your own key. If you don't have one, you can [sign up here](https://assemblyai.com/dashboard/signup) for free. Make sure that the path you pass to the `transcribe()` function corresponds to the saved filename. ```python import assemblyai as aai aai.settings.api_key = "YOUR_API_KEY" config = aai.TranscriptionConfig() transcriber = aai.Transcriber() transcript = transcriber.transcribe("wtolixa9XTg.m4a", config) print(transcript.text) ``` ## Option 2: Download video via code In this approach we download the video with a Python script instead of the command line. You can download the file with the following code: ```python import yt_dlp URLS = ['https://www.youtube.com/watch?v=wtolixa9XTg'] ydl_opts = { 'format': 'm4a/bestaudio/best', # The best audio version in m4a format 'outtmpl': '%(id)s.%(ext)s', # The output name should be the id followed by the extension 'postprocessors': [{ # Extract audio using ffmpeg 'key': 'FFmpegExtractAudio', 'preferredcodec': 'm4a', }] } with yt_dlp.YoutubeDL(ydl_opts) as ydl: error_code = ydl.download(URLS) ``` After downloading, you can use the same code from option 1 to transcribe the file: ```python import assemblyai as aai aai.settings.api_key = "YOUR_API_KEY" config = aai.TranscriptionConfig() transcriber = aai.Transcriber() transcript = transcriber.transcribe("wtolixa9XTg.m4a", config) ``` --- # Build a UI for Transcription with Gradio and Python URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/gradio-frontend Source: docs/pre-recorded-audio/guides/gradio-frontend.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Basic transcription workflows Description: Build a UI for Transcription with Gradio and Python documentation. In this guide, we'll show you how to use the [Gradio](https://www.gradio.app/) library in Python to build a simple front-end that allows you to drag-and-drop files for transcription. If you've ever gotten tired of running your Python program from the command-line, now you can have a self-hosted UI for processing your transcripts! ## Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. ## Installing Dependencies Firstly, we'll need to get all of our dependencies installed. This demo depends on both the AssemblyAI Python SDK as well as Gradio, so we'll need to run the following command to get them both installed, or updated to their latest version if you've already gotten them installed. ```bash pip install -U assemblyai gradio ``` ## Creating Your Transcription Function Now we'll move on to setting up a function that can handle transcription for us. This function will submit the filepath for the audio file you've uploaded via Gradio to our API, and will wait until the transcript has finished, or log an error to the console. ```python import assemblyai # Set your API key here, which can be found on your Dashboard as mentioned above. assemblyai.settings.api_key = "API_KEY_HERE" def transcribe(audio_path): config = assemblyai.TranscriptionConfig() transcriber = assemblyai.Transcriber() transcript = transcriber.transcribe(audio_path, config) if transcript.status == assemblyai.TranscriptStatus.error: print(f"Transcription failed: {transcript.error}") ``` ## Setting Up our Gradio UI Now we get to move on to setting up what our Gradio UI will look like. This project will be fairly simple, so it will only use a few components, but Gradio offers a wide range of components and ways to customize how they look to suit your needs. Check out their [documentation](https://www.gradio.app/docs) here for more detailed information. We'll be using Gradio's Blocks API for a more custom way of setting up our UI. Specifically, we'll use their Markdown, File, Button, and Textbox components to create a title, enable file uploads, and submit files for transcription to render them to the screen. Since we only want audio or video files to be submitted for transcription, we'll limit what can be uploaded with the `file_types` parameter. ```python import gradio with gradio.Blocks() as demo: gradio.Markdown("# AssemblyAI Transcription with Gradio") filepath = gradio.File(file_types=["audio", "video"]) transcribe_button = gradio.Button(value="Transcribe") transcript = gradio.Textbox(value="", label="Transcript") transcribe_button.click(transcribe, inputs=[filepath], outputs=[transcript]) if __name__ == "__main__": demo.launch(debug=True, show_error=True) ``` Now you can run this code to deploy your Gradio app locally, with some helpful debug information in case you add more functionality. You can also get a publicly shareable link to your app by enabling `share=True` in your `.launch()` config. Also note that you should run this in a separate file outside of this Jupyter notebook so that your program doesn't time out. We hope this helps you create better interfaces for your integrations with our API, and please reach out to support@assemblyai.com if you have any questions we can help with! --- # Detect Low Confidence Words in a Transcript URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/detecting-low-confidence-words Source: docs/pre-recorded-audio/guides/detecting-low-confidence-words.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Basic transcription workflows Description: Detect Low Confidence Words in a Transcript documentation. In this guide, we'll show you how to detect sentences that contain words with low confidence scores. Confidence scores represent how confident the model was in predicting the transcribed word. Detecting words with low confidence scores can be important for manually editing transcripts. Each transcribed word will contain a corresponding confidence score between 0.0 (low confidence) and 1.0 (high confidence). You can decide what your confidence threshold will be when implementing this logic in your application. For this guide, we will use a threshold of 0.4. ## Getting Started Before we begin, make sure you have an AssemblyAI account and an API key. You can sign up for an account and get your API key from your [dashboard](https://www.assemblyai.com/dashboard/home). This guide will use AssemblyAI's [JavaScript SDK](https://github.com/AssemblyAI/assemblyai-node-sdk). If you haven't already, install the SDK by following these [instructions](https://github.com/AssemblyAI/assemblyai-node-sdk#installation). ## Step-by-Step Instructions Import the AssemblyAI package and create an AssemblyAI object with your API key: ```javascript import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: process.env.ASSEMBLYAI_API_KEY, }); ``` Next create the transcript with your audio file, either via local audio file or URL (AssemblyAI's servers need to be able to access the URL, make sure the URL links to a downloadable file). ```javascript const transcript = await client.transcripts.transcribe({ audio_url: "./sample.mp4" }); ``` From there use the `id` from the transcript to request the transcript broken down into sentences. ```javascript let { id } = transcript; let { sentences } = await client.transcripts.sentences(id); ``` Set the confidence score threshold to a value of you choice (0.5 or less is a good start). In this guide, we'll use 0.4. ```javascript let confidenceThreshold = 0.4; ``` Next, we will filter the sentences array down to just sentences that contain words with confidence scores of under 0.4. ```javascript const sentencesWithLowConfidenceWords = (sentences, confidenceThreshold) => { return sentences.filter((sentence) => { const hasLowConfidenceWord = sentence.words.some( (word) => word.confidence < confidenceThreshold ); return hasLowConfidenceWord; }); }; const filteredSentences = sentencesWithLowConfidenceWords( sentences, confidenceThreshold ); ``` Next we'll alter the `filteredSentences` array so that the `words` array for each sentence only contains the words with confidence scores under of 0.4. ```javascript const filterScores = filteredSentences.map((item) => { return { ...item, words: item.words.filter((word) => word.confidence < confidenceThreshold), }; }); ``` Finally, we'll display the final results. The final results will include the timestamp of the sentence that contains low confidence words, the sentence, the words that scored poorly, and their scores. ```javascript expandable //This function is optional but can be used to format the timestamps from milleseconds to HH:MM:SS const formatMilliseconds = (milliseconds) => { // Calculate hours, minutes, and seconds const hours = Math.floor(milliseconds / 3600000); const minutes = Math.floor((milliseconds % 3600000) / 60000); const seconds = Math.floor((milliseconds % 60000) / 1000); // Ensure the values are displayed with leading zeros if needed const formattedHours = hours.toString().padStart(2, "0"); const formattedMinutes = minutes.toString().padStart(2, "0"); const formattedSeconds = seconds.toString().padStart(2, "0"); return `${formattedHours}:${formattedMinutes}:${formattedSeconds}`; }; //Format the final results to contain the sentence, low confidence words, timestamps, and confidence scores. const finalResults = filterScores.map((res) => { return `The following sentence at timestamp ${formatMilliseconds(res.start)} contained low confidence words: ${res.text} \n Low confidence word(s) from this sentence: ${res.words .map((res) => { return `${res.text}[score: ${res.confidence}]`; }) .join(", ")}}`; }); console.log(finalResults); ``` The output will look something like this: ``` [ 'The following sentence at timestamp 00:04:34 contained low confidence words: I am contacting you first when I could just have phoned my bank and marked you as fraud in an instant. \n' + ' Low confidence word(s) from this sentence: marked[score: 0.33049]}', 'The following sentence at timestamp 00:06:40 contained low confidence words: Sabitha, as much as I would like to help you, this is the best I can do for you. \n' + ' Low confidence word(s) from this sentence: Sabitha,[score: 0.22706]}', 'The following sentence at timestamp 00:07:37 contained low confidence words: Thank you for calling Queston. \n' + ' Low confidence word(s) from this sentence: Queston.[score: 0.16557]}' ] ``` --- # Running Bulk Transcription and Load Tests at Scale URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/bulk-transcription-and-load-tests-at-scale Source: docs/pre-recorded-audio/guides/bulk-transcription-and-load-tests-at-scale.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Batch transcription Description: Best practices for submitting large volumes of audio files for transcription, including one-off load tests. This guide applies to two closely related workloads: - **Bulk transcription.** Submitting thousands of files in a batch — for example, a nightly backfill or a one-time migration. - **Load testing.** Measuring turnaround time (TaT) and throughput before a production cutover. The guidance is shared. If you're running a load test, also read the [load test subsection](#if-youre-running-a-load-test) of *Measure and verify*. ## Key recommendations 1. **Ramp.** Submit in 15-second windows. Start at 25 requests/window and grow ~8–9% per window until you reach your target sustained rate. 2. **Measure.** Use webhooks instead of polling. Record `submit_ts`, `complete_ts`, `audio_duration`, `model`, `features`, and `status` per request. 3. **Coordinate.** Recommended for runs above 200 requests/minute, required for large bulk uploads (tens of thousands of files), and for any EU workload. The default rate limit for paid accounts is **200 parallel jobs**. If you need a higher limit, reach out to [support@assemblyai.com](mailto:support@assemblyai.com) — AssemblyAI offers custom rate limits at no additional cost. ## Before you begin **Prerequisites** - **Account balance.** If your balance hits zero mid-run, AssemblyAI drops your rate limit to 1 and invalidates your results. Bulk runs and load tests both incur standard transcription usage charges — fund your account for the full expected volume before you submit any requests. - **Rate limit.** Check your current limit on the [Rate Limits page](https://www.assemblyai.com/dashboard/home) of your dashboard and size your target submission rate against it — see [Size your target rate](#size-your-target-rate). - **Audio source.** Prefer pre-signed URLs (for example, from S3 or GCS) as your `audio_url`. Each `/v2/upload` call counts against your HTTP rate limit and adds latency proportional to file size — at bulk scale, uploads alone can exhaust your rate-limit budget. If local files are your only option, [upload them](/pre-recorded-audio/getting-started/transcribe-an-audio-file#using-the-http-api-directly) to `/v2/upload` first to get a hosted URL, and factor those uploads into your rate-limit planning. - **Pre-signed URL expiration.** Set URL TTLs long enough to outlast your expected queue time plus turnaround time. One hour is a safe default for small-to-moderate runs; use two hours or more for runs that approach your rate limit or use less-common languages. URLs that expire while a job is queued or processing surface as `4xx` errors when AssemblyAI tries to retrieve the audio. - **Completion tracking.** Configure [webhooks](/pre-recorded-audio/webhooks) (with a polling fallback) or a polling-only strategy before you start — you'll need a way to detect completion and record per-request timestamps. - **Static-IP egress.** If your audio URLs are behind a strict S3 bucket policy, contact support to enable static-IP egress for retrieval. Webhook deliveries already come from fixed IPs documented on the [whitelisting FAQ](/faq/what-ip-address-should-i-whitelist-for-assemblyai). - **Pricing reference.** See the [pricing page](https://www.assemblyai.com/pricing) for per-hour rates to plug into your cost estimate. **Workload configuration** Match these to your expected production traffic: - **Region: US or EU.** The EU region handles less traffic than US and is more sensitive to load spikes. Coordinate with our team for any EU workload, regardless of size. - **Traffic pattern.** Requests per minute at peak and whether traffic is steady or bursty. - **Models and features.** Universal-3.5 Pro, Universal-2, speaker labels, PII redaction, summarization, and so on — each has different processing characteristics, and every feature you enable adds to TaT. Audit your request body and turn off what you don't need; a faster run is also a cheaper one. - **Channels and speakers.** For call-center audio where agent and customer are on separate stereo channels, add `multichannel=true` to get per-channel utterances. For single-channel recordings with multiple speakers, use `speaker_labels=true` instead. See [Should I use Speaker Labels or Multi-channel?](/faq/should-i-use-speaker-labels-or-multi-channel) for guidance. - **Language.** Some languages, such as Hindi, Swedish, and Hebrew, scale differently and may show longer TaT. Coordinate with our team if your workload is primarily in a less common language. - **Audio format.** Format conversion and preprocessing run before transcription and contribute to overall TaT. - **Cost estimate.** Sanity-check the expected spend before you submit: total audio hours × your per-hour rate. Bulk runs are billed the same as any other transcription, so a large backfill can produce a surprisingly large invoice if you haven't projected it. ## Pilot first Before a full bulk run or load test, submit a pilot batch of 50–200 files using the exact configuration you plan to use at scale — same model, same features, same language, same webhook receiver, same error-handling logic. A pilot verifies that: - The transcripts look right, and the model and feature set match what the downstream consumer expects. - Your webhook receiver is reachable, verifying signatures, and writing results durably — or your polling loop is keeping up without hitting rate limits. - Your retry logic handles `5xx` correctly and your dead-letter path captures `4xx` without silently dropping files. - Your ramp and rate-limit controls behave as intended. The most expensive bulk-run failures almost always come from discovering a configuration mistake — a wrong model, a feature flag left off, a webhook handler that drops results silently — after the whole batch has been billed. A pilot catches these while they're cheap to fix, and gives you a realistic mean TaT to plug into [Size your target rate](#size-your-target-rate). ## When to run - **Small tests and moderate bulk runs** (well within your rate limit): US business hours (roughly 14:00–21:00 UTC) produce the most representative baseline latency, since throughput is highest during these periods. - **Large runs** (200+ requests/minute): coordinate with our team before starting. Our team will pick a window and pre-scale for you. - **EU region:** coordinate regardless of size. ## Ramp up gradually The most common mistake — for both bulk uploads and load tests — is submitting all requests at once. A gradual ramp gives the pipeline time to scale ahead of your traffic, which is what produces the lowest and most consistent turnaround times. Submitting a large spike upfront typically results in higher TaT as capacity catches up. **Recommended schedule (validated for Universal-3.5 Pro with speaker labels):** - Divide your ramp into 15-second windows. - Start at **25 requests per window**. - Grow by **~8–9% per window** until you reach your target sustained rate. - Don't pause mid-ramp. Stopping and restarting means ramping from the starting rate again, and you'll see higher latency when traffic resumes. - Do not exceed your account's rate limit during the ramp. If you're using different models or features, contact support for a tailored ramp plan — some components take longer to initialize and may need a slower ramp. To compute a ramp for any target without a lookup table, use `rate_n = ceil(25 × 1.085ⁿ)` (capped at your target rate), where `n` is the window index starting at 0. For a fully worked 400 requests/minute schedule, see [Example ramp schedule](#example-ramp-schedule) at the bottom of this page. After reaching your target rate, sustain it for at least 5–10 minutes. Your measurements are representative once p50 and p95 stay consistent over 2–3 consecutive minutes. If p50 is still falling, extend the sustain phase. ## Small runs (under 500 files) If your total run is under 500 files, you don't need the ramp. Fire a bounded thread pool at your rate limit and submit in parallel. The ramp is for sustained high rates over several minutes. ## Size your target rate Your rate limit caps how many jobs can be in progress at once, so your sustained submission rate needs to fit inside it. Pick a target rate that keeps the typical number of in-flight jobs comfortably below your limit: 1. Estimate mean turnaround time from your pilot run or the published [benchmarks](/pre-recorded-audio/benchmarks). 2. Multiply your target submission rate (requests per second) by that mean TaT to approximate the number of jobs that will be in flight at steady state. 3. Keep that number under ~80% of your rate limit. If it's higher, lower the target rate, shorten mean TaT (fewer features, shorter audio, a faster model), or [request a higher rate limit](#coordinate-with-our-team). The 20% of headroom absorbs normal variation in audio duration, warm-up effects, and webhook-receive latency. Runs sized to the limit exactly will see TaT climb as the in-flight count bumps against the cap. For planning, Universal-3.5 Pro English async typically completes a 5-minute file in about 9 seconds at p50 and 60 seconds at p95. Use that as your starting turnaround estimate before the pilot. ## Measure and verify For every run — bulk or load test — track completion and record per-request metadata. - **Use webhooks when possible.** You get a clean completion signal without polling overhead. See the [Webhooks documentation](/pre-recorded-audio/webhooks) for setup, retry behavior, and authentication. - **Run a polling fallback alongside webhooks.** Webhooks can drop for many reasons — receiver downtime, signature mismatches, transient network failures. For every submitted job, record an expected completion deadline (around 2× mean TaT from your pilot) and `GET /v2/transcript/{id}` for any job whose webhook hasn't arrived by then. The fallback protects you from silent data loss without materially increasing your rate-limit usage. - **Polling-only** every 1–2 seconds measures closer to actual completion but adds to your rate-limit budget. Use it when webhooks aren't available or when you need precise TaT during a load test. - **Retry failures** with exponential backoff for `5xx` responses — see [Implement retry server error logic](/pre-recorded-audio/guides/retry-server-error). Investigate `4xx` responses; they indicate a client-side issue that retrying won't fix. Record these fields per request: - `submit_ts` — timestamp when `POST /v2/transcript` was sent - `complete_ts` — timestamp when completion was detected - `audio_duration` — length of the audio file, in seconds - `model` — speech model used - `features` — features enabled (e.g. `speaker_labels`, `auto_highlights`, `sentiment_analysis`) - `status` — `completed` or `error` - `id` — transcript ID, for debugging with support Turnaround time = `complete_ts` − `submit_ts`. Normalize TaT by audio duration to get the real-time factor: `RTF = turnaround_time ÷ audio_duration`. An RTF of 0.5 means the API processed the file in half its audio duration. RTF is the headline metric for comparing runs across regions, models, and audio-duration buckets — raw TaT varies too much with audio length to mean anything on its own. ### Polling without exceeding the rate limit HTTP rate limits cap total API requests at **20,000 per 5 minutes** across all endpoints — submissions and polling combined. Exceeding this returns a `403` error. If webhooks aren't an option, stay within that budget. As an illustration: at a sustained 33 requests/second submission rate with ~330 jobs in flight, submissions alone consume roughly 10,000 of your budget — so poll no more often than every 15 seconds to leave headroom. Scale these numbers to your own submission rate and in-flight job count: | Polling interval | GETs/s at 330 in-flight | Total req/s (at 33 POST/s) | Within limit? | | --- | --- | --- | --- | | Every 3s | ~110 | ~143 | No | | Every 5s | ~66 | ~99 | No | | Every 10s | ~33 | ~66 | Yes | | Every 15s | ~22 | ~55 | Yes | When many jobs share the same polling interval they tend to cluster at the same second boundaries, spiking your rate-limit usage and occasionally returning `403` errors. Stagger each job's polling by ±25% of the interval (for example, `interval × random.uniform(0.75, 1.25)`) so GETs spread evenly across the window. ### If you're running a bulk job Monitor for these signals during the run: - **Healthy run:** TaT stays within ~20% of your first 5 minutes of sustained submissions, your completion queue drains steadily, and no errors arrive. - **Diagnose, don't panic:** if something goes wrong, match the signal you see to the row in [Diagnosing problems](#diagnosing-problems) and respond accordingly — each signal has a different cause and fix. ### If you're running a load test - **Separate ramp-phase from sustain-phase metrics.** Expect higher latency during ramp; use sustain-phase numbers as your benchmark. - **Report percentiles, not averages.** Track p50, p75, p90, p95, p99, and max. - **Normalize by audio duration.** Group results into duration buckets (0–5 min, 5–15 min, 15–30 min, 30–60 min) for meaningful comparison. - **Pre-scaling caveat.** If our team pre-scaled for your test, your results reflect steady-state capacity — not cold-start or scale-up behavior. ## Diagnosing problems | Signal | Meaning | Response | | --- | --- | --- | | `403` HTTP error | You've exceeded the 20,000 requests per 5 minutes HTTP rate limit (polling counts) | Increase polling interval or switch to webhooks; slow your submission rate | | `4xx` submission error (other than 403) | Client-side issue (bad request, auth, invalid audio URL, etc.) | Inspect the response body and fix the request; retrying won't help | | `5xx` submission error | Transient server-side issue | Retry with exponential backoff — see the [retry guide](/pre-recorded-audio/guides/retry-server-error) | | TaT rises with no errors | Submission rate is outpacing available capacity | Slow the ramp or extend its duration; verify you're within your rate limit | ## Reference implementation The AssemblyAI Python SDK's non-blocking `Transcriber.submit()` returns as soon as a transcript is queued, so you can drive the ramp yourself while using the SDK's `TranscriptionConfig` and exception classes. If you'd rather have the SDK handle both submission and polling for a smaller batch, see [Transcribe multiple files simultaneously](/pre-recorded-audio/guides/batch_transcription). The following script ramps submissions to approximate the recommended schedule, retries transient errors, writes unrecoverable failures to a dead-letter log, and persists submitted `file → transcript_id` pairs so the run is resumable after a crash. The table above is hand-tuned to observed pipeline behavior, so the numbers this script produces may differ by one or two requests per window. Adjust `max_rate` to your target sustained rate. ```python expandable import json import logging import math import random import time from collections import deque from pathlib import Path import assemblyai as aai from assemblyai.types import TranscriptError logger = logging.getLogger(__name__) aai.settings.api_key = "YOUR_API_KEY" # For the EU region, point the SDK at the EU base URL: # aai.settings.base_url = "https://api.eu.assemblyai.com" STATE_PATH = Path("bulk_state.jsonl") # Append-only submission log, for resume FAILURES_PATH = Path("bulk_failures.jsonl") # Dead-letter log for unrecoverable submission errors # Match your production configuration: model, features, language, etc. config = aai.TranscriptionConfig( # For split-stereo call audio (agent on one channel, customer on the other): # multichannel=True, # For single-channel audio with multiple speakers: # speaker_labels=True, ) transcriber = aai.Transcriber() def append_jsonl(path, entry): with path.open("a") as f: f.write(json.dumps(entry) + "\n") def already_submitted(path): """Return the set of file URLs successfully submitted in previous runs.""" if not path.exists(): return set() with path.open() as f: return {json.loads(line)["file"] for line in f if line.strip()} def submit_file(file_url, max_retries=3): """Submit one file without waiting for completion. Retries transient errors with exponential backoff + jitter.""" for attempt in range(max_retries + 1): try: transcript = transcriber.submit(file_url, config=config) return transcript.id except TranscriptError: if attempt == max_retries: raise time.sleep(2 ** attempt + random.random()) def submit_all(files, max_rate): """ Submit files in 15-second windows, starting at 25 requests per window and growing ~8–9% per window until reaching max_rate. Files already recorded in STATE_PATH are skipped on restart. Failed submissions are written to FAILURES_PATH for inspection and manual replay. """ done = already_submitted(STATE_PATH) remaining = deque(f for f in files if f not in done) rate = 25 # Starting requests per window window = 15 # Seconds per window growth = 1.085 # Per-window growth factor window_num = 0 while remaining: window_num += 1 batch_size = min(rate, len(remaining)) for _ in range(batch_size): file = remaining.popleft() try: transcript_id = submit_file(file) append_jsonl(STATE_PATH, {"file": file, "id": transcript_id}) except Exception as exc: logger.exception("Submission failed: %s", file) append_jsonl(FAILURES_PATH, {"file": file, "error": str(exc)}) logger.info( "Window %d | Rate: %d/window | Submitted: %d | Remaining: %d", window_num, rate, batch_size, len(remaining), ) rate = min(math.ceil(rate * growth), max_rate) if remaining: time.sleep(window) # Usage: # file_urls = ["https://example.com/audio1.mp3", ...] # submit_all(file_urls, max_rate=100) # Target: 100 requests per 15-second window # # Pick max_rate using the guidance in "Size your target rate". # Use webhooks (with a polling fallback) to track completion; account for GETs # in your rate-limit budget. ``` ## Example ramp schedule A fully worked schedule for ramping to 100 requests/window (400 requests/minute). Use it as a reference when building your own ramp, or as a starting point you can scale to a different target rate. | Time window | Requests | Cumulative | | --- | --- | --- | | 0:00–0:15 | 25 | 25 | | 0:15–0:30 | 27 | 52 | | 0:30–0:45 | 29 | 81 | | 0:45–1:00 | 31 | 112 | | 1:00–1:15 | 33 | 145 | | 1:15–1:30 | 35 | 180 | | 1:30–1:45 | 38 | 218 | | 1:45–2:00 | 41 | 259 | | 2:00–2:15 | 44 | 303 | | 2:15–2:30 | 47 | 350 | | 2:30–2:45 | 51 | 401 | | 2:45–3:00 | 55 | 456 | | 3:00–3:15 | 59 | 515 | | 3:15–3:30 | 64 | 579 | | 3:30–3:45 | 69 | 648 | | 3:45–4:00 | 75 | 723 | | 4:00–4:15 | 81 | 804 | | 4:15–4:30 | 88 | 892 | | 4:30–4:45 | 95 | 987 | | 4:45–5:00 | 100 | 1,087 | ## Coordinate with our team Reach out to [support@assemblyai.com](mailto:support@assemblyai.com) or your account manager before you submit any requests if any of these apply: - You plan to exceed **200 requests per minute**. - You're running a **large one-time upload** (tens of thousands of files or more). - You want **support available during the run** — for example, if you're running outside US business hours. - You're using the **EU region**, regardless of size. When you reach out, include: - Expected request volume and ramp schedule, broken into 15-second windows - Audio file durations and language breakdown - Speech models and features you'll enable - Whether audio is single-channel or multichannel - Preferred run window (see [When to run](#when-to-run)) AssemblyAI can pre-scale pipeline components for your traffic, raise your rate limit, and monitor the run in real time. For recurring bulk workloads (for example, nightly batch jobs), we can set up persistent scaling. ## Related pages - [Rate limits](/pre-recorded-audio/rate-limits) - [Webhooks](/pre-recorded-audio/webhooks) - [Benchmarks](/pre-recorded-audio/benchmarks) - [Implement retry server error logic](/pre-recorded-audio/guides/retry-server-error) - [Cloud endpoints and data residency](/pre-recorded-audio/select-the-region) --- # Transcribe Multiple Files Simultaneously Using the JavaScript SDK URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/sdk-node-batch Source: docs/pre-recorded-audio/guides/sdk-node-batch.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Batch transcription Description: Transcribe Multiple Files Simultaneously Using the JavaScript SDK documentation. In this guide, we'll show you how to transcribe multiple files simultaneously using the JavaScript SDK. ## Getting Started Before we begin, make sure you have an AssemblyAI account and an API key. You can sign up for an account and get your API key from your [dashboard](https://www.assemblyai.com/dashboard/home). This guide will use AssemblyAI's [JavaScript SDK](https://github.com/AssemblyAI/assemblyai-node-sdk). If you haven't already, install the SDK in your project by following these [instructions](https://github.com/AssemblyAI/assemblyai-node-sdk#installation). ## Step-by-Step Instructions Set up your application folder structure by adding an audio folder which will house the files you would like to transcribe, a transcripts folder which will house your completed transcriptions, and a new `.js` file in the root of the project. Your file structure should look like this: ``` BatchApp ├───audio │ ├───audio-1.mp3 │ └───audio-2.mp3 ├───transcripts ├───batch.js ``` In the `batch.js` file import the AssemblyAI package, as well as the node fs and node path packages. Create an AssemblyAI object with your API key: ``` import { AssemblyAI } from "assemblyai"; import * as path from 'path'; import * as fs from 'fs'; const client = new AssemblyAI({ apiKey: , }); ``` Declare the variables `audioFolder`, `files`, `filePathArr`, and `transcriptsFolder`. - `audioFolder` will be the relative path to the folder containing your audio files. - `files` will read the files in the audio folder, and return them in an array. - `filePathArr` will join the file names with the audio folder name to create the relative path to each individual file. - `transcriptsFolder` will be the relative path to the folder containing your transcription files. ``` const audioFolder = './audio'; const files = await fs.promises.readdir(audioFolder); const filePathArr = files.map(file => path.join(audioFolder, file)); const transcriptsFolder = './transcripts'; ``` Next, we'll create a promise that will submit the file path for transcription. Make sure to add the parameters for the models you would like to run. ``` const getTranscript = (filePath) => new Promise((resolve, reject) => { client.transcripts.transcribe({ audio: filePath, language_detection: true }) .then(result => resolve(result)) .catch(error => reject(error)); }); ``` Next, we will create an async function that will call the `getTranscript` function and write the transcription text from each audio file to an individual text file in the transcripts folder. ``` expandable const processFile = async (file) => { const getFileName = file.split('audio/'); //Separate the folder name and file name into substrings const fileName = getFileName[1]; //Grab the 2nd substring which is the file name const filePath = path.join(transcriptsFolder, `${fileName}.txt`); //Relative path for transcription text files. const transcript = await getTranscript(file); //Request the transcript const text = transcript.text; //Grab transcription text from the JSON response //Write the transcription text to a text file return new Promise((resolve, reject) => { fs.writeFile(filePath, text, err => { if (err) { reject(err); return; } resolve({ ok: true, message: 'Text File created!' }); }); }); } ``` Next, we will create the run function. This function will: - Create an array of unresolved promises with each promise requesting a transcript. - Use `Promise.all` to iterate over the array of unresolved promises. Then we'll call the run function ``` const run = async () => { const unresolvedPromises = filePathArr.map(processFile); await Promise.all(unresolvedPromises); } run() ``` Your final file will look like this: ``` expandable import { AssemblyAI } from "assemblyai"; import * as path from 'path'; import * as fs from 'fs'; const client = new AssemblyAI({ apiKey: , }); const audioFolder = './audio'; const files = await fs.promises.readdir(audioFolder); const filePathArr = files.map(file => path.join(audioFolder, file)); const transcriptsFolder = './transcripts' const getTranscript = (filePath) => new Promise((resolve, reject) => { client.transcripts.transcribe({ audio: filePath, language_detection: true }) .then(result => resolve(result)) .catch(error => reject(error)); }); const processFile = async (file) => { const getFileName = file.split('audio/') const fileName = getFileName[1] const filePath = path.join(transcriptsFolder, `${fileName}.txt`); const transcript = await getTranscript(file); const text = transcript.text return new Promise((resolve, reject) => { fs.writeFile(filePath, text, err => { if (err) { reject(err); return; } resolve({ ok: true, message: 'Text File created!' }); }); }); } const run = async () => { const unresolvedPromises = filePathArr.map(processFile); await Promise.all(unresolvedPromises); } run() ``` If you have any questions, please feel free to reach out to our Support team at support@assemblyai.com. --- # Transcribe Multiple Files Simultaneously Using the Python SDK URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/batch_transcription Source: docs/pre-recorded-audio/guides/batch_transcription.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Batch transcription Description: Transcribe Multiple Files Simultaneously Using the Python SDK documentation. In this guide, we'll show you how to use the AssemblyAI API to transcribe multiple audio files at once. This guide focuses on demonstrating how to use the AssemblyAI Python SDK to achieve this. ## Quickstart ```python expandable import assemblyai as aai import threading import os aai.settings.api_key = "YOUR_API_KEY" batch_folder = "audio" transcription_result_folder = "transcripts" config = aai.TranscriptionConfig() transcriber = aai.Transcriber() def transcribe_audio(audio_file): transcriber = aai.Transcriber() transcript = transcriber.transcribe(os.path.join(batch_folder, audio_file), config) if transcript.status == "completed": with open(f"{transcription_result_folder}/{audio_file}.txt", "w") as f: f.write(transcript.text) elif transcript.status == "error": print("Error: ", transcript.error) threads = [] for filename in os.listdir(batch_folder): thread = threading.Thread(target=transcribe_audio, args=(filename,)) threads.append(thread) thread.start() for thread in threads: thread.join() print("All transcriptions are complete.") ``` ## Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. ## Step-by-Step Guide Install the SDK. ```bash pip install -U assemblyai ``` Import the `assemblyai `package and set the API key. Import `threading` and `OS` Python libraries that enable concurrent task processing and file path interactions respectively. ```python import assemblyai as aai import threading import os aai.settings.api_key = "YOUR_API_KEY" ``` Set the folders. The `batch` folder contains the audio files that you want to process and transcribe. The `transcription_result_folder` stores the .txt transcript files. ```python batch_folder = "audio" transcription_result_folder = "transcripts" ``` Create a `Transcriber` object. ```python transcriber = aai.Transcriber() ``` Function to transcribe an audio file. Once the transcript is complete, a .txt file is generated to the `transcription_result_folder`. If there is an error with the transcription, it will not be processed to the results folder. ```python def transcribe_audio(audio_file): config = aai.TranscriptionConfig() transcriber = aai.Transcriber() transcript = transcriber.transcribe(os.path.join(batch_folder, audio_file), config) if transcript.status == "completed": with open(f"{transcription_result_folder}/{audio_file}.txt", "w") as f: f.write(transcript.text) elif transcript.status == "error": print("Error: ", transcript.error) ``` Open threads to transcribe each file concurrently. Once all the threads are complete you will receive the "All transcriptions are complete" message in your terminal. ```python threads = [] for filename in os.listdir(batch_folder): thread = threading.Thread(target=transcribe_audio, args=(filename,)) threads.append(thread) thread.start() for thread in threads: thread.join() print("All transcriptions are complete.") ``` ## Conclusion This guide aims to demonstrate how to use AssemblyAI Python SDK to concurrently process multiple audio files at once. The output is transcript text files for each audio file in the specified folder. Other integrations and features can be built on top of this main function. These include and are not limited to: exporting the file in different formats, adding Core Transcription or Speech Understanding features. If you have any questions, please feel free to reach out to our Support team at [support@assemblyai.com](mailto:support@assemblyai.com). --- # Transcribe from an S3 Bucket URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/transcribe_from_s3 Source: docs/pre-recorded-audio/guides/transcribe_from_s3.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Hosting audio files Description: Transcribe from an S3 Bucket documentation. AssemblyAI's Speech-to-Text APIs can be used with both local files and publicly accessible online files, but what if you want to transcribe an audio file that has restricted access? Luckily, you can do this with AssemblyAI too! Read on to learn how you can transcribe an audio file stored in an AWS S3 bucket using AssemblyAI's APIs. ## Intro In order to transcribe an audio file from an S3 bucket, AssemblyAI will need temporary access to the file. To provide this access, we'll generate a **presigned URL**, which is simply a URL that has temporary access rights baked-in. The overall process looks like this: 1. Generate a presigned URL for the S3 audio file with [boto3](https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/s3.html). 2. Pass this URL through to AssemblyAI's API with a POST request. 3. Wait until the transcription is complete, and then fetch it with a GET request. ## Prerequisites First, you'll need an AssemblyAI account. You can [sign up for a free account](https://www.assemblyai.com/dashboard/signup) if you don't already have one. Next, you'll need to **take note of your AssemblyAI API key**, which you can find on your [account dashboard](https://www.assemblyai.com/dashboard/home) after signing in. It will be on the left-hand side of the screen under _Your API Key_. You'll need the value of this key later, so leave the browser window open or copy the value into a text file. ## AWS IAM User Second, you'll need an AWS IAM user with `Programmatic` access and the `AmazonS3ReadOnlyAccess` permission. If you already have such an IAM user and you know its public and private keys, then you can move on to the next section. Otherwise, create one now as follows: First, log into AWS as a root user or as another IAM user with the appropriate access, and then go to the [IAM Management Console](https://us-east-1.console.aws.amazon.com/iamv2/home?ref=assemblyai.com#/users) to add a new user. AWS IAM Users list page with the Add users button highlighted Set the user name you would like, and select _Programmatic access_ under _Select AWS access type:_ IAM Add user step setting the user name and selecting Access key - Programmatic access Click _Next_, and then _Attach existing policies directly_. Copy and paste _AmazonS3ReadOnlyAccess_ into the _Filter policies_ search box, and then add this permission by clicking on the checkbox next to it: IAM Set permissions step with Attach existing policies directly chosen and AmazonS3ReadOnlyAccess checked Click _Next_ and add tags if you wish. Then click _Next_ and review the IAM user profile to ensure that everything looks copacetic before clicking _Create user_. IAM Add user Review screen showing user details and the AmazonS3ReadOnlyAccess policy before creating Finally, take note of the IAM user's _Access key ID_ and _Secret access key_. Again, **we will need these values later**, so copy them into a text file before moving on. **Warning** Make sure to copy the IAM user's _Secret access key_ and record it somewhere safe. Once you close the final window of the _Add user_ sequence, you will not be able to access this key again and will need to regenerate it if you forget/lose the original. ## Code First, the necessary packages are installed. ```bash pip install -U boto3 botocore ``` Then we can import them and set our relevant variable values. You'll need to edit these variables to be equivalent to the relevant values for your application: 1. `bucket_name` - The name of your AWS S3 bucket. 2. `object_name` - The name of the audio file in the S3 bucket that you want to transcribe. 3. `iam_access_id` - The access ID of the IAM user with programmatic access and S3 read permission. 4. `iam_secret_key` - The secret key of the IAM user. 5. `assembly_key` - Your AssemblyAI API key. ```python import boto3 from botocore.exceptions import ClientError import logging import requests import time bucket_name = "" object_name = "" iam_access_id = "" iam_secret_key = "" assembly_key = "" ``` From here, we simply follow the sequence outlined in the introduction of this Colab: 1. Generate a presigned URL for the S3 audio file with [boto3](https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/s3.html). ```python # Create a low-level service client with the IAM credentials. s3_client = boto3.client( "s3", aws_access_key_id=iam_access_id, aws_secret_access_key=iam_secret_key ) # Generate a pre-signed URL for the audio file that expires after 30 minutes. try: p_url = s3_client.generate_presigned_url( ClientMethod="get_object", Params={"Bucket": bucket_name, "Key": object_name}, ExpiresIn=1800, ) except ClientError as e: logging.error(e) ``` 2. Pass the presigned URL through to AssemblyAI's API with a POST request. ```python # Use your AssemblyAI API Key for authorization. headers = {"authorization": assembly_key, "content-type": "application/json"} # Specify AssemblyAI's transcription API endpoint. transcript_endpoint = "https://api.assemblyai.com/v2/transcript" # Use the presigned URL as the `audio_url` in the POST request. json = {"audio_url": p_url} # Queue the audio file for transcription with a POST request. post_response = requests.post(transcript_endpoint, json=json, headers=headers) ``` 3. Wait until the transcription is complete, and then fetch it with a GET request. ```python # Specify the endpoint of the transaction. get_endpoint = transcript_endpoint + "/" + post_response.json()["id"] # GET request the transcription. get_response = requests.get(get_endpoint, headers=headers) # If the transcription has not finished, wait util it has. while get_response.json()["status"] != "completed": get_response = requests.get(get_endpoint, headers=headers) time.sleep(5) # Once the transcription is complete, print it out. print(get_response.json()["text"]) ``` --- # Transcribe Google Drive Files URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/transcribing-google-drive-file Source: docs/pre-recorded-audio/guides/transcribing-google-drive-file.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Hosting audio files Description: Transcribe Google Drive Files documentation. ## **Step 1: Upload Your Audio File to Google Drive** - **File Requirements**: Ensure your audio file is smaller than 100MB, as files larger than this cannot be directly downloaded from Google Drive links. - **Uploading**: Log into your Google Drive account and upload the audio file you want to use. ## **Step 2: Make Your File Publicly Accessible** - **Right-Click** on the uploaded file in Google Drive. - Select **'Get Link'**. - Change the setting from “Restricted” to “Anyone with the link”. This makes the file publicly accessible. ## **Step 3: Obtain the Downloadable URL** - Click on `Copy link` to copy your shared link. - Initially, the shared link will look something like this: https://drive.google.com/file/d/1YvY3gX-4ZwY7K4r3J0THKNTvvolB3D-S/view?usp=sharing. - To make it a downloadable link, modify it to this format: `https://drive.google.com/u/0/uc?id=FILE_ID&export=download`. - **Example**: If your shared link is `https://drive.google.com/file/d/1YvY3gX-4ZwY7K4r3J0THKNTvvolB3D-S/view?usp=sharing`, change it to `https://drive.google.com/u/0/uc?id=1YvY3gX-4ZwY7K4r3J0THKNTvvolB3D-S&export=download`. Google Drive Share dialog for audio.mp3 with General access set to Anyone with the link and a Copy link button ## **Step 4: Use the URL with AssemblyAI** - Now, you can use this downloadable link in your AssemblyAI API request. This URL directly points to your audio file, allowing AssemblyAI to access and process it. ```python config = aai.TranscriptionConfig() transcriber = aai.Transcriber() audio_url = ( "https://storage.googleapis.com/aai-web-samples/5_common_sports_injuries.mp3" ) transcript = transcriber.transcribe(audio_url, config) ``` ## **Notes** - **Security**: Ensure that sharing your audio file publicly complies with your privacy and security policies. - If you prefer not to share your file publicly, you can [upload your file to our servers instead.](/api-reference/files/upload) - **File Format**: Check that your audio file is in a format supported by AssemblyAI. --- # Transcribe GitHub Files URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/transcribing-github-files Source: docs/pre-recorded-audio/guides/transcribing-github-files.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Hosting audio files Description: Transcribe GitHub Files documentation. ## Step 1: Upload Your Audio Files to a Public GitHub Repository - **File Requirements**: GitHub has a file size limit of 100MB so ensure your audio files are 100MB in size or less. The files must be in a public repository otherwise you will receive an error saying the file is not publicly accessible. For a more secure way to host files check out our [Transcribing from an S3 Bucket Cookbook](/pre-recorded-audio/guides/transcribe_from_s3). ## Step 2: Obtain the Raw Audio URL from GitHub 1. Navigate to the repository that houses the audio file. 2. Click on the audio file. On the next page, right-click the "View raw" link and select "copy the link address" from the context menu. Downloadable file URLs are formatted as `"https://github.com///raw//"` ## Step 3: Add the Audio URL to your Request POST `v2/transcript` endpoint ```JSON { "audio_url":"https://github.com/user/audio-files/raw/main/audio.mp3" } ``` Python SDK ```python config = aai.TranscriptionConfig() transcript = transcriber.transcribe("https://github.com/user/audio-files/raw/main/audio.mp3", config) ``` JavaScript SDK ```javascript const transcript = await client.transcripts.transcribe({ audio_url: "https://github.com/user/audio-files/raw/main/audio.mp3" }); ``` ## **Resources** [AssemblyAI's Supported File Types](/faq/what-audio-and-video-file-types-are-supported-by-your-api)
[Transcribe an Audio File](/pre-recorded-audio/getting-started/transcribe-an-audio-file) --- # Iterate over Speaker Labels with Make.com URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/make-speaker-labels Source: docs/pre-recorded-audio/guides/make-speaker-labels.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Speaker labels Description: Iterate over Speaker Labels with Make.com documentation. ## Introduction This is a quick guide on how to iterate over speaker labels in Make.com. This guide will return speaker labels as a readable format to a Google Doc. The end result will look like the two images below. ### Make.com Scenario: ![Image of full Make Sceanrio](/assets/img/cookbooks/makespeakers/make-scenario.png) ### Google Doc Transcript: ![Image of Final Transcript](/assets/img/cookbooks/makespeakers/make-final-transcript.png) ## Instructions ### Step 1: Transcribe the Audio Create a scenario in Make.com. Add a new module. Search for and select the AssemblyAI app and select the "Transcribe an Audio File" module. Add an audio URL. Select speaker labels and other models you’d like to run. ![Transcribe Module](/assets/img/cookbooks/makespeakers/make-transcribe-audio.png) ### Click Run once to retrieve data. ![Transcribe Module](/assets/img/cookbooks/makespeakers/make-run.png) ### Step 2: Wait for Completion Next, add the “Wait until Transcript is Ready” module. ![Wait for completion module](/assets/img/cookbooks/makespeakers/make-wait-for-completion.png) ### Select "ID" from the “Transcribe an Audio” module as input for the Transcript ID field. ![Returned data](/assets/img/cookbooks/makespeakers/make-get-id.png) ### Step 3: Get a Transcript Next, add the “Get a Transcript” module and select "ID" from the “Transcribe an Audio” module as input for the Transcript ID field. ![Get transcript module data](/assets/img/cookbooks/makespeakers/make-get-transcript.png) ### Step 4: Create a Document Search for and add the Google Docs app. From there choose the “Create a Document” module. Connect your Google account and name the Doc what you’d like. Add some filler content as well. Additionally, choose where you’d like the Doc to be located. ![Create Google doc module](/assets/img/cookbooks/makespeakers/make-create-doc.png) ### Step 5: Iterator Tool Add the Iterator tool next. The speaker label data is in the utterances array. Select that array as input from the “Transcribe an Audio File" module. This tool will be used to perform a task for each utterance in the array. The next module will repeat its action for each utterance. ![Iterator module](/assets/img/cookbooks/makespeakers/make-iterator.png) ### Step 6: Write Speaker Labels Data to Google Doc Add a module and choose the “Insert a Paragraph” module from Google Docs. Connect your Google account if it’s not already connected (you may have to reconnect it if you get a failed to load error). In the “Select a Document” drop-down, choose "by mapping". In the Document ID input field select document ID from the “Create a document” module. For appended text, you can follow the format below for a readable format. ![Insert paragrapgh module](/assets/img/cookbooks/makespeakers/make-insert-paragraph.png) ### Step 7: Run the Scenario. Run the scenario and you should get a Google Doc in your Drive with the speaker labels included. ![Image of Final Transcript](/assets/img/cookbooks/makespeakers/make-final-transcript.png) --- # Calculate the Talk / Listen Ratio of Speakers URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/talk-listen-ratio Source: docs/pre-recorded-audio/guides/talk-listen-ratio.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Speaker labels Description: Calculate the Talk / Listen Ratio of Speakers documentation. This guide will show you how to use AssemblyAI's API to calculate the talk/listen ratio of speakers in a transcript. The following code uses the [Python SDK](https://github.com/AssemblyAI/assemblyai-python-sdk). ## Quickstart ```python expandable import assemblyai as aai aai.settings.api_key = "YOUR_API_KEY" def calculate_talk_listen_ratios(transcript): """ :param transcript: AssemblyAI Transcript object :return: Dictionary with talk time, percentage, and talk-listen ratios for each speaker """ # Ensure speaker labels were enabled if not transcript.utterances: raise ValueError("Speaker labels were not enabled for this transcript.") speaker_talk_time = {} total_time = 0 for utterance in transcript.utterances: speaker = f"Speaker {utterance.speaker}" duration = utterance.end - utterance.start speaker_talk_time[speaker] = speaker_talk_time.get(speaker, 0) + duration total_time += duration # Calculate percentages and ratios result = {} for speaker, talk_time in speaker_talk_time.items(): percentage = (talk_time / total_time) * 100 result[speaker] = { "talk_time_ms": talk_time, "percentage": round(percentage, 2) } # Calculate talk-listen ratios for each speaker against all others for speaker in result.keys(): other_speakers_time = sum(talk_time for spk, talk_time in speaker_talk_time.items() if spk != speaker) if other_speakers_time > 0: ratio = speaker_talk_time[speaker] / other_speakers_time result[speaker]["talk_listen_ratio"] = round(ratio, 2) else: result[speaker]["talk_listen_ratio"] = None # Handle cases with only one speaker return result transcriber = aai.Transcriber() audio_url = ("YOUR_AUDIO_URL") config = aai.TranscriptionConfig(speaker_labels=True) transcript = transcriber.transcribe(audio_url, config) talk_listen_stats = calculate_talk_listen_ratios(transcript) print(talk_listen_stats) ``` ## Get started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up for an AssemblyAI account](https://www.assemblyai.com/dashboard/home) and get your API key from your dashboard. ## Step-by-Step Instructions Install the SDK: ```bash pip install assemblyai ``` Import the `assemblyai` package and set the API key. ```python import assemblyai as aai aai.settings.api_key = "YOUR_API_KEY" ``` Define a function called `calculate_talk_listen_ratios`, which will calculate the talk-listen ratios for all speakers from a transcript with speaker labels. Speaker labels must be enabled for the ratios to be calculated. ```python expandable def calculate_talk_listen_ratios(transcript): """ :param transcript: AssemblyAI Transcript object :return: Dictionary with talk time, percentage, and talk-listen ratios for each speaker """ # Ensure speaker labels were enabled if not transcript.utterances: raise ValueError("Speaker labels were not enabled for this transcript.") speaker_talk_time = {} total_time = 0 for utterance in transcript.utterances: speaker = f"Speaker {utterance.speaker}" duration = utterance.end - utterance.start speaker_talk_time[speaker] = speaker_talk_time.get(speaker, 0) + duration total_time += duration # Calculate percentages and ratios result = {} for speaker, talk_time in speaker_talk_time.items(): percentage = (talk_time / total_time) * 100 result[speaker] = { "talk_time_ms": talk_time, "percentage": round(percentage, 2) } # Calculate talk-listen ratios for each speaker against all others for speaker in result.keys(): other_speakers_time = sum(talk_time for spk, talk_time in speaker_talk_time.items() if spk != speaker) if other_speakers_time > 0: ratio = speaker_talk_time[speaker] / other_speakers_time result[speaker]["talk_listen_ratio"] = round(ratio, 2) else: result[speaker]["talk_listen_ratio"] = None # Handle cases with only one speaker return result ``` Define a `transcriber`, an `audio_url` set to a link to the audio file (replace the example link that is provided with your own), and a `TranscriptionConfig` with `speaker_labels=True`. Then create a transcript which will be sent to the function `calculate_talk_listen_ratios` and print out the results. ```python transcriber = aai.Transcriber() audio_url = ("https://api.assemblyai-solutions.com/storage/v1/object/public/dual-channel-phone-data/Fisher_Call_Centre/audio05851.wav") config = aai.TranscriptionConfig(speaker_labels=True) transcript = transcriber.transcribe(audio_url, config) talk_listen_stats = calculate_talk_listen_ratios(transcript) print(talk_listen_stats) ``` Example output when using the above sample audio file: ``` {'Speaker A': {'talk_time_ms': 244196, 'percentage': 42.77, 'talk_listen_ratio': 0.75}, 'Speaker B': {'talk_time_ms': 326766, 'percentage': 57.23, 'talk_listen_ratio': 1.34}} ``` --- # Plot A Speaker Timeline with Matplotlib URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/speaker_timeline Source: docs/pre-recorded-audio/guides/speaker_timeline.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Speaker labels Description: Plot A Speaker Timeline with Matplotlib documentation. In this guide, we'll show you how to plot a speaker timeline with matplotlib, using results from the speaker diarization model. ## Quickstart ```python expandable import assemblyai as aai import matplotlib.pyplot as plt aai.settings.api_key = "YOUR_API_KEY" config = aai.TranscriptionConfig(speaker_labels=True) transcriber = aai.Transcriber() transcript = transcriber.transcribe("./my-audio.mp3", config) utterances = transcript.utterances def plot_speaker_timeline(utterances): fig, ax = plt.subplots(figsize=(12, 4)) colors = ['b', 'g', 'r', 'c', 'm', 'y', 'k'] speaker_colors = {} for utterance in utterances: start = utterance.start / 60000 # in minutes end = utterance.end / 60000 # in minutes speaker = utterance.speaker if speaker not in speaker_colors: speaker_colors[speaker] = colors[len(speaker_colors) % len(colors)] # set a colour for each new speaker ax.barh(speaker, end - start, left=start, color=speaker_colors[speaker], height=0.4) # create horizontal bar plot ax.set_xlabel('Time (mins)') ax.set_ylabel('Speakers') ax.set_title('Speaker Timeline') ax.grid(True, which='both', linestyle='--', linewidth=0.5) plt.show() plot_speaker_timeline(utterances) ``` ### Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. ### Step-by-Step Instructions Install the SDK. ```bash pip install -U assemblyai !pip install -U matplotlib ``` Import the `assemblyai` package and set the API key. ```python import assemblyai as aai aai.settings.api_key = "YOUR_API_KEY" ``` Create a `TranscriptionConfig` object and set speaker labels to `True`. ```python config = aai.TranscriptionConfig(speaker_labels=True) ``` Create a `Transcriber` object. ```python transcriber = aai.Transcriber() ``` Use the Transcriber object's `transcribe` method and pass in the audio file's path and `config` object as parameters. The transcribe method saves the results of the transcription to the `Transcriber` object's `transcript` attribute. ```python transcript = transcriber.transcribe("./my-audio.mp3", config) ``` Alternatively, you can use an audio URL available on the internet. Extract the utterances from the transcript and set this to `utterances`. ```python utterances = transcript.utterances ``` Import the `matplotlib.pyplot` library. Then use the following `plot_speaker_timeline` function which results in a plot image of the speaker timeline. This function extracts the `start` and `end` timestamps of each `utterance` per `speaker` and plots the data onto the horizontal bar chart. The X and Y axis are labelled accordingly. ```python expandable import matplotlib.pyplot as plt def plot_speaker_timeline(utterances): fig, ax = plt.subplots(figsize=(12, 4)) colors = ['b', 'g', 'r', 'c', 'm', 'y', 'k'] speaker_colors = {} for utterance in utterances: start = utterance.start / 60000 # in minutes end = utterance.end / 60000 # in minutes speaker = utterance.speaker if speaker not in speaker_colors: speaker_colors[speaker] = colors[len(speaker_colors) % len(colors)] # set a colour for each new speaker ax.barh(speaker, end - start, left=start, color=speaker_colors[speaker], height=0.4) # create horizontal bar plot ax.set_xlabel('Time (mins)') ax.set_ylabel('Speakers') ax.set_title('Speaker Timeline') ax.grid(True, which='both', linestyle='--', linewidth=0.5) plt.show() ``` Finally, call the `plot_speaker_timeline` function passing `utterances` as a parameter to see the plot image result. ```python plot_speaker_timeline(utterances) ``` --- # Generate Custom Speaker Labels with Pyannote URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/Use_AssemblyAI_with_Pyannote_to_generate_custom_Speaker_Labels Source: docs/pre-recorded-audio/guides/Use_AssemblyAI_with_Pyannote_to_generate_custom_Speaker_Labels.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Speaker labels Description: Generate Custom Speaker Labels with Pyannote documentation. In this guide, we'll show you how to generate Speaker Labels using Pyannote with an AssemblyAI transcript. This can be used to generate Speaker Labels for languages we currently do not support for speaker labelling. ## Quickstart ```python expandable import os import assemblyai as aai from pyannote.audio import Pipeline import torch import pandas as pd import numpy as np # Assign your API keys HUGGING_FACE_TOKEN = os.getenv("HF_TOKEN") ASSEMBLYAI_API_KEY = os.getenv("ASSEMBLYAI_API_KEY") # Authenticate with AssemblyAI aai.settings.api_key = ASSEMBLYAI_API_KEY def transcribe_audio(audio_file, language="en"): """ Transcribe an audio file using AssemblyAI. Args: audio_file (str): Path to the audio file. language (str, optional): Language code for transcription. Defaults to "en". Returns: aai.Transcript: The transcription result. """ transcriber = aai.Transcriber(config=aai.TranscriptionConfig(language_code=language)) transcript = transcriber.transcribe(audio_file) print(f"Transcript ID: {transcript.id}") return transcript def get_speaker_labels(audio_file, transcript: aai.Transcript): """ Perform speaker diarization on an audio file and combine results with the transcript. Args: audio_file (str): Path to the audio file. transcript (aai.Transcript): The transcription result from AssemblyAI. Returns: str: A formatted string containing the transcript with speaker labels and timestamps. """ # Initialize the speaker diarization pipeline with GPU support device = torch.device("cuda" if torch.cuda.is_available() else "cpu") pipeline = Pipeline.from_pretrained( "pyannote/speaker-diarization", use_auth_token=HUGGING_FACE_TOKEN, ) if pipeline is None: raise ValueError("Failed to initialize the pipeline. Please check your authentication token and internet connection.") else: pipeline = pipeline.to(device) # Apply the pipeline to the audio file diarization = pipeline(audio_file) # Create a dictionary to store speaker segments speaker_segments = {} # Process diarization results for turn, _, speaker in diarization.itertracks(yield_label=True): start, end = turn.start, turn.end if speaker not in speaker_segments: speaker_segments[speaker] = [] speaker_segments[speaker].append((start, end)) # Convert speaker_segments to a DataFrame diarize_df = pd.DataFrame([(speaker, start, end) for speaker, segments in speaker_segments.items() for start, end in segments], columns=['speaker', 'start', 'end']) # Assign speakers to transcript words for word in transcript.words: word_start = float(word.start) / 1000 word_end = float(word.end) / 1000 overlaps = diarize_df[ (diarize_df['start'] <= word_end) & (diarize_df['end'] >= word_start) ].copy() if not overlaps.empty: overlaps['overlap'] = np.minimum(overlaps['end'], word_end) - np.maximum(overlaps['start'], word_start) word.speaker = overlaps.loc[overlaps['overlap'].idxmax(), 'speaker'] else: word.speaker = "Unknown" full_transcript = '' # Update segment speakers based on the majority speaker of its words for segment in transcript.get_sentences(): segment_start = float(segment.start) / 1000 segment_end = float(segment.end) / 1000 overlaps = diarize_df[ (diarize_df['start'] <= segment_end) & (diarize_df['end'] >= segment_start) ].copy() if not overlaps.empty: overlaps['overlap'] = np.minimum(overlaps['end'], segment_end) - np.maximum(overlaps['start'], segment_start) segment.speaker = overlaps.loc[overlaps['overlap'].idxmax(), 'speaker'] speaker_label = segment.speaker.replace('SPEAKER_', 'SPEAKER ') full_transcript += f'[{format_timestamp(segment_start)}] {speaker_label}: {segment.text}\n' else: segment.speaker = "Unknown" full_transcript += f'[{format_timestamp(segment_start)}] Unknown: {segment.text}\n' return full_transcript def format_timestamp(seconds): """ Convert seconds to a formatted timestamp string (HH:MM:SS). Args: seconds (float): Time in seconds. Returns: str: Formatted timestamp string. """ hours, remainder = divmod(int(seconds), 3600) minutes, seconds = divmod(remainder, 60) return f"{hours:02d}:{minutes:02d}:{seconds:02d}" audio_file = "audio.wav" # your local file path transcript: aai.Transcript = transcribe_audio(audio_file, language="hr") # select a language code transcript_with_speakers = get_speaker_labels(audio_file, transcript) print(transcript_with_speakers) ``` ### Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. You'll also need a HuggingFace account and API key. You can [sign up](https://huggingface.co/join) for a free account and get your API key [here](https://huggingface.co/settings/tokens). Create a **Read** type API token to ensure the necessary permissions are enabled. Browse to the [speaker-diarization](https://huggingface.co/pyannote/speaker-diarization) and [segmentation](https://huggingface.co/pyannote/segmentation) model pages and accept the **Gated Model** Terms & Conditions by entering your **Company/University**, **Website** and **Use Case** details in order to gain access to the use of these models. ### Step-by-Step Instructions Install the necessary dependencies. ```bash pip install assemblyai pyannote.audio torch pandas numpy ``` Import the necessary dependencies, assign your API keys and authenticate with AssemblyAI. ```python import os import assemblyai as aai from pyannote.audio import Pipeline import torch import pandas as pd import numpy as np # Assign your API keys HUGGING_FACE_TOKEN = os.getenv("HF_TOKEN") ASSEMBLYAI_API_KEY = os.getenv("ASSEMBLYAI_API_KEY") # Authenticate with AssemblyAI aai.settings.api_key = ASSEMBLYAI_API_KEY ``` Create the `transcribe_audio` function, this will handle the transcription process with AssemblyAI. ```python def transcribe_audio(audio_file, language="en"): """ Transcribe an audio file using AssemblyAI. Args: audio_file (str): Path to the audio file. language (str, optional): Language code for transcription. Defaults to "en". Returns: aai.Transcript: The transcription result. """ transcriber = aai.Transcriber(config=aai.TranscriptionConfig(language_code=language)) transcript = transcriber.transcribe(audio_file) print(f"Transcript ID: {transcript.id}") return transcript ``` Create the `get_speaker_labels`function, this will handle the speaker diarization model processing to generate the custom speaker labels for the transcript. Firstly, it initializes and applies the pipeline to the audio file. Secondly, it processes the diarization results and converts the speaker segments into a DataFrame so we can compare the results with the transcript. Lastly, the speaker segments are compared and assigned to the words and sentences of the transcript to create the speaker labelled transcript. ```python expandable def get_speaker_labels(audio_file, transcript: aai.Transcript): """ Perform speaker diarization on an audio file and combine results with the transcript. Args: audio_file (str): Path to the audio file. transcript (aai.Transcript): The transcription result from AssemblyAI. Returns: str: A formatted string containing the transcript with speaker labels and timestamps. """ # Initialize the speaker diarization pipeline with GPU support device = torch.device("cuda" if torch.cuda.is_available() else "cpu") pipeline = Pipeline.from_pretrained( "pyannote/speaker-diarization", use_auth_token=HUGGING_FACE_TOKEN, ) if pipeline is None: raise ValueError("Failed to initialize the pipeline. Please check your authentication token and internet connection.") else: pipeline = pipeline.to(device) # Apply the pipeline to the audio file diarization = pipeline(audio_file) # Create a dictionary to store speaker segments speaker_segments = {} # Process diarization results for turn, _, speaker in diarization.itertracks(yield_label=True): start, end = turn.start, turn.end if speaker not in speaker_segments: speaker_segments[speaker] = [] speaker_segments[speaker].append((start, end)) # Convert speaker_segments to a DataFrame diarize_df = pd.DataFrame([(speaker, start, end) for speaker, segments in speaker_segments.items() for start, end in segments], columns=['speaker', 'start', 'end']) # Assign speakers to transcript words for word in transcript.words: word_start = float(word.start) / 1000 word_end = float(word.end) / 1000 overlaps = diarize_df[ (diarize_df['start'] <= word_end) & (diarize_df['end'] >= word_start) ].copy() if not overlaps.empty: overlaps['overlap'] = np.minimum(overlaps['end'], word_end) - np.maximum(overlaps['start'], word_start) word.speaker = overlaps.loc[overlaps['overlap'].idxmax(), 'speaker'] else: word.speaker = "Unknown" full_transcript = '' # Update segment speakers based on the majority speaker of its words for segment in transcript.get_sentences(): segment_start = float(segment.start) / 1000 segment_end = float(segment.end) / 1000 overlaps = diarize_df[ (diarize_df['start'] <= segment_end) & (diarize_df['end'] >= segment_start) ].copy() if not overlaps.empty: overlaps['overlap'] = np.minimum(overlaps['end'], segment_end) - np.maximum(overlaps['start'], segment_start) segment.speaker = overlaps.loc[overlaps['overlap'].idxmax(), 'speaker'] speaker_label = segment.speaker.replace('SPEAKER_', 'SPEAKER ') full_transcript += f'[{format_timestamp(segment_start)}] {speaker_label}: {segment.text}\n' else: segment.speaker = "Unknown" full_transcript += f'[{format_timestamp(segment_start)}] Unknown: {segment.text}\n' return full_transcript ``` If you know the number of speakers in advance, you can use the `num_speakers` parameter to set the number of speakers: ```python # Apply the pipeline to the audio file diarization = pipeline(audio_file, num_speakers=4) ``` You can also provide upper/lower bands on the number of speakers using the `min_speakers` and `max_speakers` parameters: ```python # Apply the pipeline to the audio file diarization = pipeline(audio_file, min_speakers=2, max_speakers=5) ``` Create the `format_timestamp`, this will handle the timestamps conversion to improve the readability of the final speaker labelled transcript. ```python def format_timestamp(seconds): """ Convert seconds to a formatted timestamp string (HH:MM:SS). Args: seconds (float): Time in seconds. Returns: str: Formatted timestamp string. """ hours, remainder = divmod(int(seconds), 3600) minutes, seconds = divmod(remainder, 60) return f"{hours:02d}:{minutes:02d}:{seconds:02d}" ``` Finally, select a local file and call the functions to generate and print your custom Speaker Labelled transcript. ```python audio_file = "audio.wav" # your local file path transcript: aai.Transcript = transcribe_audio(audio_file, language="hr") # select a language code transcript_with_speakers = get_speaker_labels(audio_file, transcript) print(transcript_with_speakers) ``` Here's an example speaker labelled output from a Croatian file: ``` [00:00:05] SPEAKER 04: Nalazimo se u Centro Zagreba, u parku Zrinjevac, gdje je kao što vidite jako ljepo, vreme je prekrasno, a danas ćemo ljude pitati što im se sviđa u Zagrebu ili što im se možda ne sviđa u Zagrebu. [00:00:42] SPEAKER 04: Dobar dan, može jednokratko pitanje samo. [00:00:46] SPEAKER 04: Može? [00:00:48] SPEAKER 04: Evo lako, što vam se najviše sviđa u Zagrebu? [00:00:50] SPEAKER 07: Što mi se najviše sviđa u Zagrebu? [00:00:53] SPEAKER 07: E sad, teško pitanje, ali trenutno mi se najviše sviđa što nije klasična jesen, nego više prođeče u zraku. [00:01:06] SPEAKER 07: Dobre. [00:01:09] SPEAKER 07: Može sigurnost još uvijek s osišam sigurno u Zagrebu. [00:01:13] SPEAKER 04: I po noći? [00:01:15] SPEAKER 07: Pa po noći ne šetam baš toliko po noći, ali centar grada mi je dosta siguran, osvijetljen i to mi je okej. ``` --- # Use Speaker Diarization with Async Chunking URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/speaker-diarization-with-async-chunking Source: docs/pre-recorded-audio/guides/speaker-diarization-with-async-chunking.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Speaker labels Description: Use Speaker Diarization with Async Chunking documentation. This guide uses AssemblyAI and [Nvidia's NeMo](https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/titanet_large) framework. We'll be using TitaNet, a state of the art open source model that is trained for speaker recognition tasks. TitaNet will allow us to generate audio embeddings for speakers, which can be used to identify semantic similarity matches between two speakers. ## Quickstart ```python expandable import assemblyai as aai import requests import json import time import requests import copy from pydub import AudioSegment import os import nemo.collections.asr as nemo_asr from pydub import AudioSegment speaker_model = nemo_asr.models.EncDecSpeakerLabelModel.from_pretrained("nvidia/speakerverification_en_titanet_large") assemblyai_key = "YOUR_API_KEY" headers = { "authorization": assemblyai_key } def get_transcript(transcript_id): polling_endpoint = f"https://api.assemblyai.com/v2/transcript/{transcript_id}" while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': # print("Transcript ID:", transcript_id) return(transcription_result) break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) def download_wav(presigned_url, output_filename): # Download the WAV file from the presigned URL response = requests.get(presigned_url) if response.status_code == 200: print("downloading...") with open(output_filename, 'wb') as f: f.write(response.content) print("successfully downloaded file:", output_filename) else: raise Exception("Failed to download file, status code: {}".format(response.status_code)) # Function to identify the longest monologue of each speaker from each clip # you pass in the utterances and it returns the longest monologue from each speaker on that file def find_longest_monologues(utterances): longest_monologues = {} current_monologue = {} last_speaker = None # Track the last speaker to identify interruptions for utterance in utterances: speaker = utterance['speaker'] start_time = utterance['start'] end_time = utterance['end'] if speaker not in current_monologue: current_monologue[speaker] = {"start": start_time, "end": end_time} longest_monologues[speaker] = [] else: # Extend monologue only if it's the same speaker speaking continuously if current_monologue[speaker]["end"] == start_time and last_speaker == speaker: current_monologue[speaker]["end"] = end_time else: monologue_length = current_monologue[speaker]["end"] - current_monologue[speaker]["start"] new_entry = (monologue_length, copy.deepcopy(current_monologue[speaker])) if len(longest_monologues[speaker]) < 1 or monologue_length > min(longest_monologues[speaker], key=lambda x: x[0])[0]: if len(longest_monologues[speaker]) == 1: longest_monologues[speaker].remove(min(longest_monologues[speaker], key=lambda x: x[0])) longest_monologues[speaker].append(new_entry) current_monologue[speaker] = {"start": start_time, "end": end_time} last_speaker = speaker # Update the last speaker # Check the last monologue for each speaker for speaker, monologue in current_monologue.items(): monologue_length = monologue["end"] - monologue["start"] new_entry = (monologue_length, monologue) if len(longest_monologues[speaker]) < 1 or monologue_length > min(longest_monologues[speaker], key=lambda x: x[0])[0]: if len(longest_monologues[speaker]) == 1: longest_monologues[speaker].remove(min(longest_monologues[speaker], key=lambda x: x[0])) longest_monologues[speaker].append(new_entry) return longest_monologues # Create clips of each long monologue and embed the clip # you pass in the file path and the longest monologue objects returned by the find_longest_monologues function. # This function will create new audio file clips which contain only the longest monologue from each speaker def clip_and_store_utterances(audio_file, longest_monologues): # Load the full conversation audio full_audio = AudioSegment.from_wav(audio_file) full_audio = full_audio.set_channels(1) utterance_clips = [] for speaker, monologues in longest_monologues.items(): for _, monologue in monologues: start_ms = monologue['start'] end_ms = monologue['end'] clip = full_audio[start_ms:end_ms] clip_filename = f"{speaker}_monologue_{start_ms}_{end_ms}.wav" clip.export(clip_filename, format="wav") utterance_clips.append({ 'clip_filename': clip_filename, 'start': start_ms, 'end': end_ms, 'speaker': speaker }) print("Total Number of Monologue Clips Found: ", len(utterance_clips)) return utterance_clips # This function uses NeMO to compare two files def compare_embeddings(utterance_clip, reference_file): verification_result = speaker_model.verify_speakers( utterance_clip, reference_file ) return verification_result file_one = "YOUR_FILE_1" file_two = "YOUR_FILE_2" file_three = "YOUR_FILE_3" download_wav(file_one, "testone.wav") download_wav(file_two, "testtwo.wav") download_wav(file_three, "testthree.wav") # Store utterances from each clip, keyed by clip index clip_utterances = {} # Dictionary to track known speaker identities across all clips # Maps current clip speaker labels to a unified speaker label speaker_identity_map = {} def process_clips(clip_transcript_ids, audio_files): global clip_utterances, speaker_identity_map # This will store the longest clip filenames for each speaker from the previous clips previous_speaker_clips = {} for clip_index, (transcript_id, audio_file) in enumerate(zip(clip_transcript_ids, audio_files)): transcript = get_transcript(transcript_id) utterances = transcript['utterances'] clip_utterances[clip_index] = utterances # Store utterances for the current clip longest_monologues = find_longest_monologues(utterances) # Process the longest monologues for clipping and storing current_speaker_clips = {} for speaker, monologue_data in longest_monologues.items(): clip_and_store_utterances(audio_file, {speaker: monologue_data}) longest_clip = f"{speaker}_monologue_{monologue_data[0][1]['start']}_{monologue_data[0][1]['end']}.wav" current_speaker_clips[speaker] = longest_clip if clip_index == 0: speaker_identity_map = {speaker: speaker for speaker in longest_monologues.keys()} previous_speaker_clips = current_speaker_clips.copy() else: # Compare all new speakers against all base speakers from previous clips for new_speaker, new_clip in current_speaker_clips.items(): for base_speaker, base_clip in previous_speaker_clips.items(): if compare_embeddings(new_clip, base_clip): speaker_identity_map[new_speaker] = base_speaker break else: # If no match is found, assign a new label new_label = chr(ord(max(speaker_identity_map.values(), key=lambda x: ord(x))) + 1) speaker_identity_map[new_speaker] = new_label # Update the previous_speaker_clips for the next iteration previous_speaker_clips.update(current_speaker_clips) # Update utterances with the new speaker labels for the current clip for utterance in clip_utterances[clip_index]: original_speaker = utterance['speaker'] # Update only if there's a change in speaker identity if original_speaker in speaker_identity_map: utterance['speaker'] = speaker_identity_map[original_speaker] # Add your clip transcript IDs clip_transcript_ids = [ "YOUR_TRANSCRIPT_ID_1", "YOUR_TRANSCRIPT_ID_2", "YOUR_TRANSCRIPT_ID_3" ] # Add filepaths to your downloaded files audio_files = [ "/testone.wav", "/testtwo.wav", "/testthree.wav" ] process_clips(clip_transcript_ids, audio_files) def display_transcript(transcript_data): for clip_index, utterances in transcript_data.items(): print(f"Clip {clip_index + 1}:") for utterance in utterances: speaker = utterance['speaker'] text = utterance['text'] print(f" Speaker {speaker}: {text}") print("\n") # Add an extra newline for spacing between display_transcript(clip_utterances) ``` ## Get Started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up for an AssemblyAI account](https://www.assemblyai.com/dashboard/home) and get your API key from your dashboard. ## Step-by-step instructions ### Install Dependencies ```bash pip install pytorch pip install nemo_toolkit['all'] pip install ffmpeg pip install assemblyai ``` ### AssemblyAI Setup, Transcript Setup, and Load the Model Using NeMO In this section, we'll import dependencies and add functions to transcribe and store transcript IDs if needed. ```python import assemblyai as aai import requests import json import time import requests import copy from pydub import AudioSegment import os import nemo.collections.asr as nemo_asr speaker_model = nemo_asr.models.EncDecSpeakerLabelModel.from_pretrained("nvidia/speakerverification_en_titanet_large") assemblyai_key = "YOUR_API_KEY" headers = { "authorization": assemblyai_key } ``` ### Helper Functions The function below requests a transcript based on a transcript ID. ```python def get_transcript(transcript_id): polling_endpoint = f"https://api.assemblyai.com/v2/transcript/{transcript_id}" while True: transcription_result = requests.get(polling_endpoint, headers=headers).json() if transcription_result['status'] == 'completed': # print("Transcript ID:", transcript_id) return(transcription_result) break elif transcription_result['status'] == 'error': raise RuntimeError(f"Transcription failed: {transcription_result['error']}") else: time.sleep(3) ``` Our main inference function will make use of these functions: ```python expandable def download_wav(presigned_url, output_filename): # Download the WAV file from the presigned URL response = requests.get(presigned_url) if response.status_code == 200: print("downloading...") with open(output_filename, 'wb') as f: f.write(response.content) print("successfully downloaded file:", output_filename) else: raise Exception("Failed to download file, status code: {}".format(response.status_code)) # Function to identify the longest monologue of each speaker from each clip # you pass in the utterances and it returns the longest monologue from each speaker on that file def find_longest_monologues(utterances): longest_monologues = {} current_monologue = {} last_speaker = None # Track the last speaker to identify interruptions for utterance in utterances: speaker = utterance['speaker'] start_time = utterance['start'] end_time = utterance['end'] if speaker not in current_monologue: current_monologue[speaker] = {"start": start_time, "end": end_time} longest_monologues[speaker] = [] else: # Extend monologue only if it's the same speaker speaking continuously if current_monologue[speaker]["end"] == start_time and last_speaker == speaker: current_monologue[speaker]["end"] = end_time else: monologue_length = current_monologue[speaker]["end"] - current_monologue[speaker]["start"] new_entry = (monologue_length, copy.deepcopy(current_monologue[speaker])) if len(longest_monologues[speaker]) < 1 or monologue_length > min(longest_monologues[speaker], key=lambda x: x[0])[0]: if len(longest_monologues[speaker]) == 1: longest_monologues[speaker].remove(min(longest_monologues[speaker], key=lambda x: x[0])) longest_monologues[speaker].append(new_entry) current_monologue[speaker] = {"start": start_time, "end": end_time} last_speaker = speaker # Update the last speaker # Check the last monologue for each speaker for speaker, monologue in current_monologue.items(): monologue_length = monologue["end"] - monologue["start"] new_entry = (monologue_length, monologue) if len(longest_monologues[speaker]) < 1 or monologue_length > min(longest_monologues[speaker], key=lambda x: x[0])[0]: if len(longest_monologues[speaker]) == 1: longest_monologues[speaker].remove(min(longest_monologues[speaker], key=lambda x: x[0])) longest_monologues[speaker].append(new_entry) return longest_monologues # Create clips of each long monologue and embed the clip # you pass in the file path and the longest monologue objects returned by the find_longest_monologues function. # This function will create new audio file clips which contain only the longest monologue from each speaker def clip_and_store_utterances(audio_file, longest_monologues): # Load the full conversation audio full_audio = AudioSegment.from_wav(audio_file) full_audio = full_audio.set_channels(1) utterance_clips = [] for speaker, monologues in longest_monologues.items(): for _, monologue in monologues: start_ms = monologue['start'] end_ms = monologue['end'] clip = full_audio[start_ms:end_ms] clip_filename = f"{speaker}_monologue_{start_ms}_{end_ms}.wav" clip.export(clip_filename, format="wav") utterance_clips.append({ 'clip_filename': clip_filename, 'start': start_ms, 'end': end_ms, 'speaker': speaker }) print("Total Number of Monologue Clips Found: ", len(utterance_clips)) return utterance_clips # This function uses NeMO to compare two files def compare_embeddings(utterance_clip, reference_file): verification_result = speaker_model.verify_speakers( utterance_clip, reference_file ) return verification_result ``` ### Inference Add the links to the WAV file clips you have stored on your server. ```python file_one = "YOUR_FILE_1" file_two = "YOUR_FILE_2" file_three = "YOUR_FILE_3" download_wav(file_one, "testone.wav") download_wav(file_two, "testtwo.wav") download_wav(file_three, "testthree.wav") ``` ```python expandable from pydub import AudioSegment # Store utterances from each clip, keyed by clip index clip_utterances = {} # Dictionary to track known speaker identities across all clips # Maps current clip speaker labels to a unified speaker label speaker_identity_map = {} def process_clips(clip_transcript_ids, audio_files): global clip_utterances, speaker_identity_map # This will store the longest clip filenames for each speaker from the previous clips previous_speaker_clips = {} for clip_index, (transcript_id, audio_file) in enumerate(zip(clip_transcript_ids, audio_files)): transcript = get_transcript(transcript_id) utterances = transcript['utterances'] clip_utterances[clip_index] = utterances # Store utterances for the current clip longest_monologues = find_longest_monologues(utterances) # Process the longest monologues for clipping and storing current_speaker_clips = {} for speaker, monologue_data in longest_monologues.items(): clip_and_store_utterances(audio_file, {speaker: monologue_data}) longest_clip = f"{speaker}_monologue_{monologue_data[0][1]['start']}_{monologue_data[0][1]['end']}.wav" current_speaker_clips[speaker] = longest_clip if clip_index == 0: speaker_identity_map = {speaker: speaker for speaker in longest_monologues.keys()} previous_speaker_clips = current_speaker_clips.copy() else: # Compare all new speakers against all base speakers from previous clips for new_speaker, new_clip in current_speaker_clips.items(): for base_speaker, base_clip in previous_speaker_clips.items(): if compare_embeddings(new_clip, base_clip): speaker_identity_map[new_speaker] = base_speaker break else: # If no match is found, assign a new label new_label = chr(ord(max(speaker_identity_map.values(), key=lambda x: ord(x))) + 1) speaker_identity_map[new_speaker] = new_label # Update the previous_speaker_clips for the next iteration previous_speaker_clips.update(current_speaker_clips) # Update utterances with the new speaker labels for the current clip for utterance in clip_utterances[clip_index]: original_speaker = utterance['speaker'] # Update only if there's a change in speaker identity if original_speaker in speaker_identity_map: utterance['speaker'] = speaker_identity_map[original_speaker] # Add your clip transcript IDs clip_transcript_ids = [ "YOUR_TRANSCRIPT_ID_1", "YOUR_TRANSCRIPT_ID_2", "YOUR_TRANSCRIPT_ID_3" ] # Add filepaths to your downloaded files audio_files = [ "/testone.wav", "/testtwo.wav", "/testthree.wav" ] process_clips(clip_transcript_ids, audio_files) ``` ### New Clip Utterances The clip utterances returned by the `process_clips` function will contain the corrected utterances, which can be seen by printing out the utterances or by using the `display_transcript` function. ```python print(clip_utterances) ``` ```python def display_transcript(transcript_data): for clip_index, utterances in transcript_data.items(): print(f"Clip {clip_index + 1}:") for utterance in utterances: speaker = utterance['speaker'] text = utterance['text'] print(f" Speaker {speaker}: {text}") print("\n") # Add an extra newline for spacing between display_transcript(clip_utterances) ``` ## Additional Resources - [Async chunking](https://github.com/AssemblyAI-Solutions/async-chunking) - [AsyncChunkPy: Near-Realtime Python Speech-to-Text App](https://github.com/AssemblyAI-Solutions/async-chunk-py) - [Guide for Identifying speakers across multiple podcasts with AssemblyAI and TitaNet](https://docs.google.com/document/d/1xdOvY1LM2lGUNRZCf3kNQLSs5OCUfo74_MHDGm_jYc0/edit?tab=t.0#heading=h.59f375jz72wm) - [Multi-Speaker Voice Identification and Diarization with AssemblyAI, Pinecone, and Nvidia's TitaNet Model](https://colab.research.google.com/drive/1vqpYcPLEjDjMJ9WvP8C1gvOIflBHs0VS?usp=sharing) --- # Setup A Speaker Identification System using Pinecone & Nvidia TitaNet URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/titanet-speaker-identification Source: docs/pre-recorded-audio/guides/titanet-speaker-identification.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Speaker labels Description: Setup A Speaker Identification System using Pinecone & Nvidia TitaNet documentation. This guide will demonstrate how to build an advanced speaker recognition and diarization system that you can use to identify speakers across multiple audio files. It will use: - AssemblyAI for transcription and initial diarization. - Nvidia's TitaNet model for speaker embedding generation. - Pinecone for efficient similarity search of speaker embeddings. ## Quickstart ```python expandable from pinecone import Pinecone, ServerlessSpec import assemblyai as aai import requests import os from pydub import AudioSegment import mimetypes import wave from nemo.collections.asr.models import EncDecSpeakerLabelModel import torch import numpy as np import uuid from sklearn.metrics.pairwise import cosine_similarity # Obtain from your Pinecone dashboard. pc = Pinecone(api_key="PINECONE_KEY_HERE") # Obtain from your AssemblyAI dashboard. aai.settings.api_key = "AAI_KEY_HERE" def transcribe(file_url): config = aai.TranscriptionConfig(speaker_labels=True) # Speaker labels must be enabled for this Cookbook. transcriber = aai.Transcriber(config=config) transcript = transcriber.transcribe(file_url) return transcript.json_response def download_and_convert_to_wav(url, output_dir="./content/converted_audio"): # Create the output directory if it doesn't exist. os.makedirs(output_dir, exist_ok=True) # Extract filename from URL. filename = url.split("/")[-1].split("?")[0] base_filename, file_extension = os.path.splitext(filename) # Download the file. response = requests.get(url) if response.status_code == 200: # Determine the file type. content_type = response.headers.get("content-type") if content_type: guessed_extension = mimetypes.guess_extension(content_type) if guessed_extension: file_extension = guessed_extension # Save the downloaded file. downloaded_file = os.path.join(output_dir, filename) with open(downloaded_file, "wb") as f: f.write(response.content) # Generate the WAV file name. wav_filename = f"{base_filename}.wav" wav_file = os.path.join(output_dir, wav_filename) # Load the audio file. audio = AudioSegment.from_file(downloaded_file) # Convert to mono if it's stereo. if audio.channels > 1: print("Setting channels to 1.") audio = audio.set_channels(1) # Export as WAV. audio.export(wav_file, format="wav") print(f"File converted and saved as: {wav_file}") # Remove the original downloaded file if it's different from the WAV file. if downloaded_file != wav_file: os.remove(downloaded_file) # Ensure the WAV file is single channel. with wave.open(wav_file, "rb") as wf: n_channels = wf.getnchannels() if n_channels > 1: print(f"Converting {n_channels} channels to mono...") # Read the frames. frames = wf.readframes(wf.getnframes()) # Get other parameters. params = wf.getparams() # Close the file. wf.close() # Convert to mono. mono_frames = b"".join([frames[i::n_channels] for i in range(n_channels)]) # Write the mono WAV file. with wave.open(wav_file, "wb") as wf: wf.setparams( (1, params.sampwidth, params.framerate, params.nframes, params.comptype, params.compname) ) wf.writeframes(mono_frames) print("Conversion to mono complete.") return wav_file else: print(f"Failed to download the file. Status code: {response.status_code}") return None def add_speaker_embedding_to_pinecone(speaker_name, speaker_embedding, unique_id=None): # Ensure the embedding is a 1D numpy array. if isinstance(speaker_embedding, torch.Tensor): embedding_np = speaker_embedding.squeeze().cpu().numpy() elif isinstance(speaker_embedding, np.ndarray): embedding_np = speaker_embedding.squeeze() else: raise ValueError("Unsupported embedding type. Expected torch.Tensor or numpy.ndarray") # Ensure the embedding is the correct shape if embedding_np.shape != (192,): raise ValueError(f"Expected embedding of shape (192,), but got {embedding_np.shape}") # Convert to list for Pinecone embedding_list = embedding_np.tolist() # Generate a unique ID if not provided if unique_id is None: unique_id = f"speaker_{speaker_name}_{uuid.uuid4().hex[:8]}" # Create the metadata dictionary metadata = {"speaker_name": speaker_name} # Upsert the vector to Pinecone upsert_response = index.upsert(vectors=[(unique_id, embedding_list, metadata)]) print(f"Upserted embedding for speaker {speaker_name} with ID {unique_id}") return unique_id def find_closest_speaker(utterance_embedding, local_embeddings=None, local_only=False, threshold=0.5): def cosine_sim(a, b): return cosine_similarity(a.reshape(1, -1), b.reshape(1, -1))[0][0] best_match = {"speaker_name": "No match found", "score": 0} # Local embeddings processing. if local_embeddings is not None: for speaker_name, embedding in local_embeddings.items(): score = cosine_sim(utterance_embedding, embedding) if score > best_match["score"]: print("Identified speaker " + speaker_name + " confidence " + str(score)) best_match = {"speaker_name": speaker_name, "score": score} # Pinecone query (if not local_only and local_embeddings is empty or not provided) if not local_only and (local_embeddings is None or len(local_embeddings) == 0): results = index.query(vector=utterance_embedding.tolist(), top_k=1, include_metadata=True) if results["matches"]: pinecone_match = results["matches"][0] pinecone_score = pinecone_match["score"] if pinecone_score > best_match["score"]: best_match = {"speaker_name": pinecone_match["metadata"]["speaker_name"], "score": pinecone_score} # Check if the best match meets the threshold. if best_match["score"] < threshold: return "No match found", 0 return best_match["speaker_name"], best_match["score"] def identify_speakers_from_utterances(transcript, wav_file, min_utterance_length=5000, match_all_utterances=False): utterances = transcript["utterances"] speaker_model = EncDecSpeakerLabelModel.from_pretrained("nvidia/speakerverification_en_titanet_large") known_speakers = {} unknown_speakers = {} unknown_count = 0 unknown_folder = "unknown_speaker_utterances" os.makedirs(unknown_folder, exist_ok=True) audio_file_name = os.path.basename(wav_file) full_audio = AudioSegment.from_wav(wav_file) def get_suitable_utterance(speaker, min_length): suitable_utterances = [ u for u in utterances if u["speaker"] == speaker and (u["end"] - u["start"]) >= min_length ] if suitable_utterances: return max(suitable_utterances, key=lambda u: u["end"] - u["start"]) return max((u for u in utterances if u["speaker"] == speaker), key=lambda u: u["end"] - u["start"]) # First pass: Identify speakers. for speaker in set(u["speaker"] for u in utterances): if speaker not in known_speakers and speaker not in unknown_speakers: suitable_utterance = get_suitable_utterance(speaker, min_utterance_length) start_ms = suitable_utterance["start"] end_ms = suitable_utterance["end"] utterance_audio = full_audio[start_ms:end_ms] temp_wav = "temp_utterance.wav" utterance_audio.export(temp_wav, format="wav") embedding = speaker_model.get_embedding(temp_wav) os.remove(temp_wav) speaker_name, score = find_closest_speaker(embedding) print(f"Speaker: {speaker}, Closest match: {speaker_name}, Score: {score}") if score > 0.5: # Adjust threshold as needed. known_speakers[speaker] = speaker_name print(f"Identified as known speaker: {speaker}") else: unknown_count += 1 unknown_name = f"Unknown Speaker {chr(64 + unknown_count)}" unknown_wav = f"{unknown_folder}/unknown_speaker_{unknown_count}_from_{audio_file_name}" utterance_audio.export(unknown_wav, format="wav") unknown_speakers[speaker] = { "name": unknown_name, "wav_file": unknown_wav, "duration": end_ms - start_ms, } print(f"New unknown speaker detected: {unknown_name} (Duration: {(end_ms - start_ms)/1000:.2f}s)") # Second pass: Replace speaker names. for utterance in utterances: if utterance["speaker"] in known_speakers: utterance["speaker"] in known_speakers[utterance["speaker"]] elif utterance["speaker"] in unknown_speakers: utterance["speaker"] = unknown_speakers[utterance["speaker"]]["name"] # Third pass: Match all utterances if requested. if match_all_utterances: print("Matching all utterances individually...") for utterance in utterances: start_ms = utterance["start"] end_ms = utterance["end"] utterance_audio = full_audio[start_ms:end_ms] temp_wav = "temp_utterance.wav" utterance_audio.export(temp_wav, format="wav") embedding = speaker_model.get_embedding(temp_wav) os.remove(temp_wav) new_speaker_name, score = find_closest_speaker(embedding) if score > 0.5 and new_speaker_name != utterance["speaker"]: print(f"Speaker change detected: '{utterance['speaker']}' -> '{new_speaker_name}' (Score: {score})") print(f"Utterance: {utterance['text'][:50]}...") utterance["speaker"] = new_speaker_name return utterances, unknown_speakers pc.create_index( name="speaker-embeddings", dimension=192, # Replace with model-specific dimensions - 192 is for TitaNet-Large. metric="cosine", # Replace with your model metric. spec=ServerlessSpec( cloud="aws", region="us-east-1", ), ) # Connect to our new index. index = pc.Index("speaker-embeddings") speaker_model = EncDecSpeakerLabelModel.from_pretrained("nvidia/speakerverification_en_titanet_large") elon_fingerprint = download_and_convert_to_wav( "https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/musk_fingerprinting.wav?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L211c2tfZmluZ2VycHJpbnRpbmcud2F2IiwiaWF0IjoxNzAwNDM4OTAwLCJleHAiOjE3MzE5NzQ5MDB9.O3QOJSBqFNb1sg4nurwSFA13xIPHyKuon3UfHcFYit0&t=2023-11-20T00%3A08%3A20.071Z" ) altman_fingerprint = download_and_convert_to_wav( "https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/sam_altman_fingerprint.mp3?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L3NhbV9hbHRtYW5fZmluZ2VycHJpbnQubXAzIiwiaWF0IjoxNzAwNjY5NjI0LCJleHAiOjE3MzIyMDU2MjR9._1yuMGzBhFcHr7xv76160Hb_SC-mH_Wv3_qX-S7XsTU&t=2023-11-22T16%3A13%3A44.103Z" ) lex_fingerprint = download_and_convert_to_wav( "https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/lex_fridman_fingerprint.mp3?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L2xleF9mcmlkbWFuX2ZpbmdlcnByaW50Lm1wMyIsImlhdCI6MTcwMDY2OTc4OSwiZXhwIjoxNzMyMjA1Nzg5fQ.PUWTOLIHl4dcrWjh2ZJ_2TBaxMXpcU-x6OcvUDe6ZXQ&t=2023-11-22T16%3A16%3A29.697Z" ) known_speakers = {"Elon Musk": elon_fingerprint, "Sam Altman": altman_fingerprint, "Lex Fridman": lex_fingerprint} # Upload the known speakers. for speaker, audio_file in known_speakers.items(): print("***") print(speaker) print(audio_file) embedding = speaker_model.get_embedding(audio_file) add_speaker_embedding_to_pinecone(speaker, embedding) audio_file = download_and_convert_to_wav( "https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/musk_fingerprinting.wav?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L211c2tfZmluZ2VycHJpbnRpbmcud2F2IiwiaWF0IjoxNzAwNDM4OTAwLCJleHAiOjE3MzE5NzQ5MDB9.O3QOJSBqFNb1sg4nurwSFA13xIPHyKuon3UfHcFYit0&t=2023-11-20T00%3A08%3A20.071Z" ) utterance_embedding = speaker_model.get_embedding(audio_file) results = index.query(vector=utterance_embedding.tolist(), top_k=3, include_metadata=True) print(results) # Example: Conversation Between Sam Altman and Elon Musk transcript_obj = transcribe("https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/elon_altman_interview_clipped.wav?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L2Vsb25fYWx0bWFuX2ludGVydmlld19jbGlwcGVkLndhdiIsImlhdCI6MTcwMDY5MDU3OSwiZXhwIjoxNzMyMjI2NTc5fQ.4qZHvVRGhNGttfcpcfXDcJkJe_tbkc_2Bvs4i51SNSE&t=2023-11-22T22%3A02%3A59.429Z") wav_file = download_and_convert_to_wav("https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/elon_altman_interview_clipped.wav?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L2Vsb25fYWx0bWFuX2ludGVydmlld19jbGlwcGVkLndhdiIsImlhdCI6MTcwMDY5MDU3OSwiZXhwIjoxNzMyMjI2NTc5fQ.4qZHvVRGhNGttfcpcfXDcJkJe_tbkc_2Bvs4i51SNSE&t=2023-11-22T22%3A02%3A59.429Z") identified_utterances, unknown_speakers = identify_speakers_from_utterances(transcript_obj, wav_file) for utterance in identified_utterances: print(f"{utterance['speaker']}: {utterance['text']}") ``` ## Initial Setup First, you'll need to [sign up for an AssemblyAI account](https://www.assemblyai.com/dashboard/signup) and obtain your API key from your [account dashboard](https://www.assemblyai.com/dashboard/home). Then, [sign up for a Pinecone account](https://app.pinecone.io/?sessionType=signup) and obtain your API key from "API Keys" on the sidebar of your dashboard. Also note that any files you use for this Cookbook should be in WAV format. While not a requirement for AssemblyAI, TitaNet requires WAV format. ## Installing Dependencies Now we'll need to install the necessary libraries and frameworks for this project. Please note that this process can take several minutes to complete. ```bash pip install -U Cython torch nemo_toolkit ffmpeg pydub pinecone-client assemblyai hydra-core pytorch_lightning huggingface_hub==0.23.5 librosa transformers pandas inflect webdataset sentencepiece youtokentome pyannote-audio editdistance jiwer lhotse datasets ``` ## Pinecone Setup In this section, we'll import Pinecone, create a new index for our speaker embeddings, and connect to the index. Please enter your Pinecone API key in the placeholder below. ```python from pinecone import Pinecone, ServerlessSpec # Obtain from your Pinecone dashboard. pc = Pinecone(api_key="PINECONE_KEY_HERE") pc.create_index( name="speaker-embeddings", dimension=192, # Replace with model-specific dimensions - 192 is for TitaNet-Large. metric="cosine", # Replace with your model metric. spec=ServerlessSpec( cloud="aws", region="us-east-1", ), ) # Connect to our new index. index = pc.Index("speaker-embeddings") ``` ## AssemblyAI Setup Now we'll set up AssemblyAI for transcription and diarization. We'll import the necessary modules and create a function to transcribe our audio files with speaker labels enabled. Please enter your AssemblyAI API key in the cell below. ```python import assemblyai as aai aai.settings.api_key = "AAI_KEY_HERE" def transcribe(file_url): config = aai.TranscriptionConfig(speaker_labels=True) # Speaker labels must be enabled for this Cookbook. transcriber = aai.Transcriber(config=config) transcript = transcriber.transcribe(file_url) return transcript.json_response ``` We'll also need to create a `download_and_convert_to_wav` helper function. This function allows us to take file URLs, download them, then convert them to WAV format. If the URLs are already in WAV format, then they're just downloaded. The files must be in WAV format to work properly with the TitaNet. ```python expandable import requests import os from pydub import AudioSegment import mimetypes import wave def download_and_convert_to_wav(url, output_dir="./content/converted_audio"): # Create the output directory if it doesn't exist. os.makedirs(output_dir, exist_ok=True) # Extract filename from URL. filename = url.split("/")[-1].split("?")[0] base_filename, file_extension = os.path.splitext(filename) # Download the file. response = requests.get(url) if response.status_code == 200: # Determine the file type. content_type = response.headers.get("content-type") if content_type: guessed_extension = mimetypes.guess_extension(content_type) if guessed_extension: file_extension = guessed_extension # Save the downloaded file. downloaded_file = os.path.join(output_dir, filename) with open(downloaded_file, "wb") as f: f.write(response.content) # Generate the WAV file name. wav_filename = f"{base_filename}.wav" wav_file = os.path.join(output_dir, wav_filename) # Load the audio file. audio = AudioSegment.from_file(downloaded_file) # Convert to mono if it's stereo. if audio.channels > 1: print("Setting channels to 1.") audio = audio.set_channels(1) # Export as WAV. audio.export(wav_file, format="wav") print(f"File converted and saved as: {wav_file}") # Remove the original downloaded file if it's different from the WAV file. if downloaded_file != wav_file: os.remove(downloaded_file) # Ensure the WAV file is single channel. with wave.open(wav_file, "rb") as wf: n_channels = wf.getnchannels() if n_channels > 1: print(f"Converting {n_channels} channels to mono...") # Read the frames. frames = wf.readframes(wf.getnframes()) # Get other parameters. params = wf.getparams() # Close the file. wf.close() # Convert to mono. mono_frames = b"".join([frames[i::n_channels] for i in range(n_channels)]) # Write the mono WAV file. with wave.open(wav_file, "wb") as wf: wf.setparams( (1, params.sampwidth, params.framerate, params.nframes, params.comptype, params.compname) ) wf.writeframes(mono_frames) print("Conversion to mono complete.") return wav_file else: print(f"Failed to download the file. Status code: {response.status_code}") return None ``` ## NVIDIA's TitaNet Model Setup Next we'll import `torch` and `nemo`, then connect to and load [NVIDIA's TitaNet model](https://huggingface.co/nvidia/speakerverification_en_titanet_large). This model allows us to generate speaker embeddings to create speaker fingerprints. It also enables the conversion of utterances into embeddings for comparison with our fingerprints. ```python from nemo.collections.asr.models import EncDecSpeakerLabelModel speaker_model = EncDecSpeakerLabelModel.from_pretrained("nvidia/speakerverification_en_titanet_large") ``` We'll now define an `add_speaker_embedding_to_pinecone` function to add our speaker embeddings to the Pinecone database. ```python expandable import torch import numpy as np import uuid def add_speaker_embedding_to_pinecone(speaker_name, speaker_embedding, unique_id=None): # Ensure the embedding is a 1D numpy array. if isinstance(speaker_embedding, torch.Tensor): embedding_np = speaker_embedding.squeeze().cpu().numpy() elif isinstance(speaker_embedding, np.ndarray): embedding_np = speaker_embedding.squeeze() else: raise ValueError("Unsupported embedding type. Expected torch.Tensor or numpy.ndarray") # Ensure the embedding is the correct shape if embedding_np.shape != (192,): raise ValueError(f"Expected embedding of shape (192,), but got {embedding_np.shape}") # Convert to list for Pinecone embedding_list = embedding_np.tolist() # Generate a unique ID if not provided if unique_id is None: unique_id = f"speaker_{speaker_name}_{uuid.uuid4().hex[:8]}" # Create the metadata dictionary metadata = {"speaker_name": speaker_name} # Upsert the vector to Pinecone upsert_response = index.upsert(vectors=[(unique_id, embedding_list, metadata)]) print(f"Upserted embedding for speaker {speaker_name} with ID {unique_id}") return unique_id ``` ## Add Thumbprints to our Pinecone Database Below we'll use chunks of the speakers' conversations to generate speaker embeddings and add them to our vector database. Later on, we'll show how to take an audio file with speakers not in the vector database and obtain the data required to generate new speaker fingerprints to be uploaded to the Pinecone database. ```python elon_fingerprint = download_and_convert_to_wav( "https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/musk_fingerprinting.wav?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L211c2tfZmluZ2VycHJpbnRpbmcud2F2IiwiaWF0IjoxNzAwNDM4OTAwLCJleHAiOjE3MzE5NzQ5MDB9.O3QOJSBqFNb1sg4nurwSFA13xIPHyKuon3UfHcFYit0&t=2023-11-20T00%3A08%3A20.071Z" ) altman_fingerprint = download_and_convert_to_wav( "https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/sam_altman_fingerprint.mp3?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L3NhbV9hbHRtYW5fZmluZ2VycHJpbnQubXAzIiwiaWF0IjoxNzAwNjY5NjI0LCJleHAiOjE3MzIyMDU2MjR9._1yuMGzBhFcHr7xv76160Hb_SC-mH_Wv3_qX-S7XsTU&t=2023-11-22T16%3A13%3A44.103Z" ) lex_fingerprint = download_and_convert_to_wav( "https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/lex_fridman_fingerprint.mp3?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L2xleF9mcmlkbWFuX2ZpbmdlcnByaW50Lm1wMyIsImlhdCI6MTcwMDY2OTc4OSwiZXhwIjoxNzMyMjA1Nzg5fQ.PUWTOLIHl4dcrWjh2ZJ_2TBaxMXpcU-x6OcvUDe6ZXQ&t=2023-11-22T16%3A16%3A29.697Z" ) known_speakers = {"Elon Musk": elon_fingerprint, "Sam Altman": altman_fingerprint, "Lex Fridman": lex_fingerprint} # Upload the known speakers. for speaker, audio_file in known_speakers.items(): print("***") print(speaker) print(audio_file) embedding = speaker_model.get_embedding(audio_file) add_speaker_embedding_to_pinecone(speaker, embedding) ``` Now we can query our Pinecone database to ensure that our embeddings were uploaded successfully. ```python audio_file = download_and_convert_to_wav( "https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/musk_fingerprinting.wav?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L211c2tfZmluZ2VycHJpbnRpbmcud2F2IiwiaWF0IjoxNzAwNDM4OTAwLCJleHAiOjE3MzE5NzQ5MDB9.O3QOJSBqFNb1sg4nurwSFA13xIPHyKuon3UfHcFYit0&t=2023-11-20T00%3A08%3A20.071Z" ) utterance_embedding = speaker_model.get_embedding(audio_file) results = index.query(vector=utterance_embedding.tolist(), top_k=3, include_metadata=True) print(results) ``` ## Creating Functions to Find the Closest Speaker and Identify Speakers of Utterances ### Speaker Identification Function The `find_closest_speaker` function is a crucial component of our speaker identification system. It compares a given utterance embedding to known speaker embeddings and identifies the closest match. ```python expandable from sklearn.metrics.pairwise import cosine_similarity def find_closest_speaker(utterance_embedding, local_embeddings=None, local_only=False, threshold=0.5): def cosine_sim(a, b): return cosine_similarity(a.reshape(1, -1), b.reshape(1, -1))[0][0] best_match = {"speaker_name": "No match found", "score": 0} # Local embeddings processing. if local_embeddings is not None: for speaker_name, embedding in local_embeddings.items(): score = cosine_sim(utterance_embedding, embedding) if score > best_match["score"]: print("Identified speaker " + speaker_name + " confidence " + str(score)) best_match = {"speaker_name": speaker_name, "score": score} # Pinecone query (if not local_only and local_embeddings is empty or not provided) if not local_only and (local_embeddings is None or len(local_embeddings) == 0): results = index.query(vector=utterance_embedding.tolist(), top_k=1, include_metadata=True) if results["matches"]: pinecone_match = results["matches"][0] pinecone_score = pinecone_match["score"] if pinecone_score > best_match["score"]: best_match = {"speaker_name": pinecone_match["metadata"]["speaker_name"], "score": pinecone_score} # Check if the best match meets the threshold. if best_match["score"] < threshold: return "No match found", 0 return best_match["speaker_name"], best_match["score"] ``` ### Speaker Identification from Utterances The `identify_speakers_from_utterances` function is the core of our speaker identification system. It processes a transcript with utterances and identifies speakers, handling both known and unknown voices. ```python expandable def identify_speakers_from_utterances(transcript, wav_file, min_utterance_length=5000, match_all_utterances=False): utterances = transcript["utterances"] speaker_model = EncDecSpeakerLabelModel.from_pretrained("nvidia/speakerverification_en_titanet_large") known_speakers = {} unknown_speakers = {} unknown_count = 0 unknown_folder = "unknown_speaker_utterances" os.makedirs(unknown_folder, exist_ok=True) audio_file_name = os.path.basename(wav_file) full_audio = AudioSegment.from_wav(wav_file) def get_suitable_utterance(speaker, min_length): suitable_utterances = [ u for u in utterances if u["speaker"] == speaker and (u["end"] - u["start"]) >= min_length ] if suitable_utterances: return max(suitable_utterances, key=lambda u: u["end"] - u["start"]) return max((u for u in utterances if u["speaker"] == speaker), key=lambda u: u["end"] - u["start"]) # First pass: Identify speakers. for speaker in set(u["speaker"] for u in utterances): if speaker not in known_speakers and speaker not in unknown_speakers: suitable_utterance = get_suitable_utterance(speaker, min_utterance_length) start_ms = suitable_utterance["start"] end_ms = suitable_utterance["end"] utterance_audio = full_audio[start_ms:end_ms] temp_wav = "temp_utterance.wav" utterance_audio.export(temp_wav, format="wav") embedding = speaker_model.get_embedding(temp_wav) os.remove(temp_wav) speaker_name, score = find_closest_speaker(embedding) print(f"Speaker: {speaker}, Closest match: {speaker_name}, Score: {score}") if score > 0.5: # Adjust threshold as needed. known_speakers[speaker] = speaker_name print(f"Identified as known speaker: {speaker}") else: unknown_count += 1 unknown_name = f"Unknown Speaker {chr(64 + unknown_count)}" unknown_wav = f"{unknown_folder}/unknown_speaker_{unknown_count}_from_{audio_file_name}" utterance_audio.export(unknown_wav, format="wav") unknown_speakers[speaker] = { "name": unknown_name, "wav_file": unknown_wav, "duration": end_ms - start_ms, } print(f"New unknown speaker detected: {unknown_name} (Duration: {(end_ms - start_ms)/1000:.2f}s)") # Second pass: Replace speaker names. for utterance in utterances: if utterance["speaker"] in known_speakers: utterance["speaker"] in known_speakers[utterance["speaker"]] elif utterance["speaker"] in unknown_speakers: utterance["speaker"] = unknown_speakers[utterance["speaker"]]["name"] # Third pass: Match all utterances if requested. if match_all_utterances: print("Matching all utterances individually...") for utterance in utterances: start_ms = utterance["start"] end_ms = utterance["end"] utterance_audio = full_audio[start_ms:end_ms] temp_wav = "temp_utterance.wav" utterance_audio.export(temp_wav, format="wav") embedding = speaker_model.get_embedding(temp_wav) os.remove(temp_wav) new_speaker_name, score = find_closest_speaker(embedding) if score > 0.5 and new_speaker_name != utterance["speaker"]: print(f"Speaker change detected: '{utterance['speaker']}' -> '{new_speaker_name}' (Score: {score})") print(f"Utterance: {utterance['text'][:50]}...") utterance["speaker"] = new_speaker_name return utterances, unknown_speakers ``` ## Examples: Speaker Identification and Diarization To demonstrate the capabilities of our speaker identification and diarization system, we'll cover several examples. We'll start with a straightforward case and progressively move to more complex scenarios. ### Example 1: Conversation Between Sam Altman and Elon Musk Our first example is a simple conversation between two well-known figures: Elon Musk and Sam Altman. This example will showcase how our system performs with clear, distinct voices in a controlled setting. ```python transcript_obj = transcribe("https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/elon_altman_interview_clipped.wav?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L2Vsb25fYWx0bWFuX2ludGVydmlld19jbGlwcGVkLndhdiIsImlhdCI6MTcwMDY5MDU3OSwiZXhwIjoxNzMyMjI2NTc5fQ.4qZHvVRGhNGttfcpcfXDcJkJe_tbkc_2Bvs4i51SNSE&t=2023-11-22T22%3A02%3A59.429Z") wav_file = download_and_convert_to_wav("https://api.assemblyai-solutions.com/storage/v1/object/sign/sam_training_bucket/elon_altman_interview_clipped.wav?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1cmwiOiJzYW1fdHJhaW5pbmdfYnVja2V0L2Vsb25fYWx0bWFuX2ludGVydmlld19jbGlwcGVkLndhdiIsImlhdCI6MTcwMDY5MDU3OSwiZXhwIjoxNzMyMjI2NTc5fQ.4qZHvVRGhNGttfcpcfXDcJkJe_tbkc_2Bvs4i51SNSE&t=2023-11-22T22%3A02%3A59.429Z") identified_utterances, unknown_speakers = identify_speakers_from_utterances(transcript_obj, wav_file) for utterance in identified_utterances: print(f"{utterance['speaker']}: {utterance['text']}") ``` --- # Use Automatic Language Detection as a Separate Step From Transcription URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/automatic-language-detection-separate Source: docs/pre-recorded-audio/guides/automatic-language-detection-separate.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Automatic Language Detection Description: Use Automatic Language Detection as a Separate Step From Transcription documentation. In this guide, we'll show you how to perform automatic language detection separately from the transcription process. For the transcription, the file then gets then routed to either our [Universal-3.5 Pro or Universal-2](/pre-recorded-audio/select-the-speech-model) model class, depending on the supported language. This workflow is designed to be cost-effective, slicing the first 60 seconds of audio and running it through Universal-2 ALD, which detects 99 languages, at a cost of $0.002 per transcript for this language detection workflow (not including the total transcription cost). ## Get started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://assemblyai.com/dashboard/signup) for a free account and get your API key from your dashboard. ## Step-by-step instructions Install the SDK: ```bash pip install assemblyai ``` Import the `assemblyai` package and set your API key: ```python import assemblyai as aai aai.settings.api_key = "YOUR_API_KEY" ``` Create a set with all supported languages for _Universal_. You can find them in our [documentation here](/pre-recorded-audio/supported-languages). ```python expandable supported_languages_for_universal = { "en", "en_au", "en_uk", "en_us", "es", "fr", "de", "it", "pt", "nl", "hi", "ja", "zh", "fi", "ko", "pl", "ru", "tr", "uk", "vi", } ``` Define a `Transcriber`. Note that here we don't pass in a global `TranscriptionConfig`, but later apply different ones during the `transcribe()` call. ```python transcriber = aai.Transcriber() ``` Define two helper functions: - `detect_language()` performs language detection on the [first 60 seconds](/api-reference/transcripts/submit#request.body.audio_end_at) of the audio and returns the language code. - `transcribe_file()` performs the transcription. For this, the identified language is applied and either Universal-3.5 Pro or Universal-2 is used depending on the supported language. ```python def detect_language(audio_url): config = aai.TranscriptionConfig( audio_end_at=60000, # first 60 seconds (in milliseconds) language_detection=True, speech_models=["universal-2"], ) transcript = transcriber.transcribe(audio_url, config=config) return transcript.json_response["language_code"] def transcribe_file(audio_url, language_code): config = aai.TranscriptionConfig( language_code=language_code, speech_models=( ["universal-3-5-pro", "universal-2"] if language_code in supported_languages_for_universal else ["universal-2"] ), ) transcript = transcriber.transcribe(audio_url, config=config) return transcript ``` Test the code with different audio files. For each file, we apply both helper functions sequentially to first identify the language and then transcribe the file. ```python audio_urls = [ "https://storage.googleapis.com/aai-web-samples/public_benchmarking_portugese.mp3", "https://storage.googleapis.com/aai-web-samples/public_benchmarking_spanish.mp3", "https://storage.googleapis.com/aai-web-samples/slovenian_luka_doncic_interview.mp3", "https://storage.googleapis.com/aai-web-samples/5_common_sports_injuries.mp3", ] for audio_url in audio_urls: language_code = detect_language(audio_url) print("Identified language:", language_code) transcript = transcribe_file(audio_url, language_code) print("Transcript:", transcript.text[:100], "...") ``` --- # Route to Default Language if Language Confidence is Low URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/automatic-language-detection-route-default-language Source: docs/pre-recorded-audio/guides/automatic-language-detection-route-default-language.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Automatic Language Detection Description: Route to Default Language if Language Confidence is Low documentation. This guide will show you how to use AssemblyAI's API to resubmit a request using a default language if the Automatic Language Detection's `language_confidence` is below a certain threshold. ## Getting started Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up for an AssemblyAI account](https://www.assemblyai.com/dashboard/home) and get your API key from your dashboard. ## Step-by-step instructions Install the SDK: ```bash npm install assemblyai ``` ```bash pip install assemblyai ``` Import the `assemblyai` package and set the API key. ```js import { AssemblyAI } from "assemblyai"; const client = new AssemblyAI({ apiKey: "YOUR_API_KEY", }); ``` ```python import assemblyai as aai aai.settings.api_key = "YOUR_API_KEY" ``` Define a `default_language`, which should be set to the [language code](/pre-recorded-audio/supported-languages) that will be used to rerun the transcript if language detection runs with low `language_confidence`. ```js const default_language = "LANGUAGE_CODE"; ``` ```python default_language = "LANGUAGE_CODE" ``` Define an `audio_url` that is set to a link to the audio file. Define and set the parameters `audio: audioUrl` and `language_detection: true`. We also need to define our `language_confidence_threshold`. For the purposes of this example, we'll set it to 0.8, representing 80% confidence. If a transcript ends up with a `language_confidence` below this value, the transcript will error out and will return the transcript using the `default_language`. ```js const audioUrl = "https://example.org/audio.mp3"; const params = { audio: audioUrl, language_detection: true, language_confidence_threshold: 0.8, speech_models: ["universal-3-5-pro", "universal-2"], // Add any other params }; ``` ```python transcriber = aai.Transcriber() audio_url = ("https://example.org/audio.mp3") config = aai.TranscriptionConfig(language_detection=True, language_confidence_threshold=0.8, speech_models=["universal-3-5-pro", "universal-2"]) transcript = transcriber.transcribe(audio_url, config) ``` You can handle the error safely by checking the error message and rerunning the transcript with the `language_code` set to the `default_language`. The error handling flow works as follows: 1. If there is no error, the transcript ID and text are printed 2. If there is an error: - Check if it's related to `language_confidence` being below threshold - If so: - Print message about rerunning with default language - Create new transcript with `default_language` as the `language_code` - Print new transcript ID and text - If not: - Print the error message When rerunning with the default language, the configuration is updated to: - Turn off `language_detection` - Remove `language_confidence_threshold` - Set `language_code` to the `default_language` You will not be charged for the first transcript if there is an error. You will only be charged for the transcript that processes successfully. ```js expandable const run = async (params) => { const transcript = await client.transcripts.transcribe(params); if (transcript.status === "error") { if ( transcript.error.includes( "below the requested confidence threshold value" ) ) { console.log( `${transcript.error}. Running transcript again with language set to '${default_language}'.` ); params = { ...params, language_detection: false, language_confidence_threshold: null, language_code: default_language, }; run(params); return; } console.log(transcript.error); return; } console.log(`Transcript ID: ${transcript.id}`); console.log(transcript.text); }; run(params); ``` ```python if transcript.error: if "below the requested confidence threshold value" in transcript.error: print(f"{transcript.error}. Running transcript again with language set to '{default_language}'.") new_config = aai.TranscriptionConfig(language_code=default_language, speech_models=["universal-3-5-pro", "universal-2"]) transcript = transcriber.transcribe(audio_url, new_config) print(f"Transcript ID: {transcript.id}") print(transcript.text) else: print(transcript.error) else: print(f"Transcript ID: {transcript.id}") print(transcript.text) ``` --- # Create Custom Length Subtitles URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/subtitle_creation_by_word_count Source: docs/pre-recorded-audio/guides/subtitle_creation_by_word_count.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Subtitles Description: Create Custom Length Subtitles documentation. While our SRT/VTT endpoints do allow you to customize the maximum number of characters per caption using the chars_per_caption URL parameter in your API requests, there are some use-cases that require a custom number of words in each subtitle. In this guide, we will demonstrate how to construct these subtitles yourself in Python! ## Quickstart ```python expandable import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" config = aai.TranscriptionConfig() transcriber = aai.Transcriber() transcript = transcriber.transcribe("./my-audio.mp3", config) def second_to_timecode(x: float) -> str: hour, x = divmod(x, 3600) minute, x = divmod(x, 60) second, x = divmod(x, 1) millisecond = int(x * 1000.) return '%.2d:%.2d:%.2d,%.3d' % (hour, minute, second, millisecond) def generate_subtitles_by_word_count(transcript, words_per_line): output = [] subtitle_index = 1 # Start subtitle index at 1 word_count = 0 current_words = [] for sentence in transcript.get_sentences(): for word in sentence.words: current_words.append(word) word_count += 1 if word_count >= words_per_line or word == sentence.words[-1]: start_time = second_to_timecode(current_words[0].start / 1000) end_time = second_to_timecode(current_words[-1].end / 1000) subtitle_text = " ".join([word.text for word in current_words]) output.append(str(subtitle_index)) output.append("%s --> %s" % (start_time, end_time)) output.append(subtitle_text) output.append("") current_words = [] # Reset for the next subtitle word_count = 0 # Reset word count subtitle_index += 1 return output subs = generate_subtitles_by_word_count(transcript, 6) with open(f"{transcript.id}.srt", 'w') as o: final = '\n'.join(subs) o.write(final) print("SRT file generated.") ``` ## Step-by-Step Instructions ```bash pip install -U assemblyai ``` Create a `main.py` file and import the `assemblyai` package and set the API key. ```python import assemblyai as aai aai.settings.api_key = "YOUR-API-KEY" ``` Create a Transcriber object. ```python config = aai.TranscriptionConfig() transcriber = aai.Transcriber() ``` Use the Transcriber object's transcribe method and pass in the audio file's path as a parameter. The transcribe method saves the results of the transcription to the Transcriber object's transcript attribute. ```python transcript = transcriber.transcribe("./my-audio.mp3", config) ``` Alternatively, you can pass in the URL of the publicly accessible audio file on the internet. ```python transcript = transcriber.transcribe("https://storage.googleapis.com/aai-docs-samples/espn.m4a", config) ``` Define a function that converts seconds to timecodes ```python def second_to_timecode(x: float) -> str: hour, x = divmod(x, 3600) minute, x = divmod(x, 60) second, x = divmod(x, 1) millisecond = int(x * 1000.) return '%.2d:%.2d:%.2d,%.3d' % (hour, minute, second, millisecond) ``` Define a function that iterates through the transcripts object to construct a list according to the number of words per subtitle ```python expandable def generate_subtitles_by_word_count(transcript, words_per_line): output = [] subtitle_index = 1 # Start subtitle index at 1 word_count = 0 current_words = [] for sentence in transcript.get_sentences(): for word in sentence.words: current_words.append(word) word_count += 1 if word_count >= words_per_line or word == sentence.words[-1]: start_time = second_to_timecode(current_words[0].start / 1000) end_time = second_to_timecode(current_words[-1].end / 1000) subtitle_text = " ".join([word.text for word in current_words]) output.append(str(subtitle_index)) output.append("%s --> %s" % (start_time, end_time)) output.append(subtitle_text) output.append("") current_words = [] # Reset for the next subtitle word_count = 0 # Reset word count subtitle_index += 1 return output ``` Generate your subtitle file ```python subs = generate_subtitles_by_word_count(transcript, 6) with open(f"{transcript.id}.srt", 'w') as o: final = '\n'.join(subs) o.write(final) print("SRT file generated.") ``` Run your script. ```bash python main.py ``` --- # Create Subtitles with Speaker Labels URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/speaker_labelled_subtitles Source: docs/pre-recorded-audio/guides/speaker_labelled_subtitles.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Subtitles Description: Create Subtitles with Speaker Labels documentation. ## Quickstart ```python expandable import assemblyai as aai # SETTINGS aai.settings.api_key = "YOUR-API-KEY" filename = "YOUR-FILE-NAME" transcriber = aai.Transcriber(config=aai.TranscriptionConfig(speaker_labels=True)) transcript = transcriber.transcribe(filename) # Maximum number of words per subtitle max_words_per_subtitle = 6 # Color assignments for speakers speaker_colors = { "A": "red", "B": "orange", "C": "yellow", "D": "yellowgreen", "E": "green", "F": "lightskyblue", "G": "purple", "H": "mediumpurple", "I": "pink", "J": "brown", } # Process transcription segments def process_segments(segments): srt_content = "" subtitle_index = 1 for segment in segments: speaker = segment.speaker color = speaker_colors.get(speaker, "black") # Default color is black # Split text into words and group into chunks words = segment.words for i in range(0, len(words), max_words_per_subtitle): chunk = words[i:i + max_words_per_subtitle] start_time = chunk[0].start # -1 indicates continuation end_time = chunk[-1].end srt_content += create_subtitle(subtitle_index, start_time, end_time, chunk, color) subtitle_index += 1 return srt_content # Create a single subtitle def create_subtitle(index, start_time, end_time, words, color): text = "" for word in words: text += word.text + ' ' start_srt = format_time(start_time) end_srt = format_time(end_time) return f"{index}\n{start_srt} --> {end_srt}\n{text}\n\n" # Format time in SRT style def format_time(milliseconds): hours, remainder = divmod(milliseconds, 3600000) minutes, remainder = divmod(remainder, 60000) seconds, milliseconds = divmod(remainder, 1000) return f"{int(hours):02}:{int(minutes):02}:{int(seconds):02},{int(milliseconds):03}" # Generate SRT content sentences = transcript.get_sentences() srt_content = process_segments(sentences) # Save to SRT file with open(filename + '.srt', 'w') as file: file.write(srt_content) print(f"SRT file generated: {filename}.srt") ``` This Colab will demonstrate how to use AssemblyAI's [Speaker Diarization](/pre-recorded-audio/label-speakers) model together to format subtitles according to their respective speaker. ## Step-by-step guide Before we begin, make sure you have an AssemblyAI account and an API key. You can [sign up](https://www.assemblyai.com/dashboard/signup) for an AssemblyAI account and get your API key from your [dashboard](https://www.assemblyai.com/dashboard/home). ```bash pip install assemblyai ``` First, we will configure our API key as well as our file to be transcribed. Then, we decide on a number of words we want to have per subtitle. Lastly, we transcribe our file. ```python import assemblyai as aai # SETTINGS aai.settings.api_key = "YOUR-API-KEY" filename = "YOUR-FILE-NAME" transcriber = aai.Transcriber(config=aai.TranscriptionConfig(speaker_labels=True)) transcript = transcriber.transcribe(filename) # Maximum number of words per subtitle max_words_per_subtitle = 6 ``` ## How the code works `speaker_colors` is a dictionary that maps speaker identifiers (like "A", "B", "C", etc.) to specific colors. Each speaker in the transcription will be associated with a unique color in the subtitles. When Speaker Diarization is enabled, sentences in our API response have a speaker code under the `speaker` key. We use the speaker code to determine the color of the subtitle text. ```python expandable # Color assignments for speakers speaker_colors = { "A": "red", "B": "orange", "C": "yellow", "D": "yellowgreen", "E": "green", "F": "lightskyblue", "G": "purple", "H": "mediumpurple", "I": "pink", "J": "brown", } # Process transcription segments def process_segments(segments): srt_content = "" subtitle_index = 1 for segment in segments: speaker = segment.speaker color = speaker_colors.get(speaker, "black") # Default color is black # Split text into words and group into chunks words = segment.words for i in range(0, len(words), max_words_per_subtitle): chunk = words[i:i + max_words_per_subtitle] start_time = chunk[0].start # -1 indicates continuation end_time = chunk[-1].end srt_content += create_subtitle(subtitle_index, start_time, end_time, chunk, color) subtitle_index += 1 return srt_content # Create a single subtitle def create_subtitle(index, start_time, end_time, words, color): text = "" for word in words: text += word.text + ' ' start_srt = format_time(start_time) end_srt = format_time(end_time) return f"{index}\n{start_srt} --> {end_srt}\n{text}\n\n" # Format time in SRT style def format_time(milliseconds): hours, remainder = divmod(milliseconds, 3600000) minutes, remainder = divmod(remainder, 60000) seconds, milliseconds = divmod(remainder, 1000) return f"{int(hours):02}:{int(minutes):02}:{int(seconds):02},{int(milliseconds):03}" ``` Our last step is to generate and save our subtitle file! ```python # Generate SRT content sentences = transcript.get_sentences() srt_content = process_segments(sentences) # Save to SRT file with open(filename + '.srt', 'w') as file: file.write(srt_content) print(f"SRT file generated: {filename}.srt") ``` --- # Generate Subtitles for Videos URL: https://www.assemblyai.com/docs/pre-recorded-audio/guides/subtitles Source: docs/pre-recorded-audio/guides/subtitles.mdx Navigation: Pre-recorded STT > Guides > Tutorials > Subtitles Description: Generate Subtitles for Videos documentation. You can export your completed transcripts in SRT or VTT format, which can be used for subtitles and closed captions in videos. Once your transcript status shows as completed, you can make a request to the appropriate endpoint to export your transcript in SRT or VTT format. In this Colab, we'll walk through the process of generating subtitles for videos using the AssemblyAI API. ## SRT and VTT Subtitles Formats ### How do SRT Files Work? SRT (SubRip Text) files are commonly used to store subtitles for videos. The format is plain text, and it contains the timing information for each subtitle along with the subtitle text itself. Here's a breakdown of how the format works: - Each subtitle entry consists of an index number, start time, end time, and text. - The index number is a sequential number starting from 1. - The start and end times are given in the format `hours:minutes:seconds,milliseconds` and are separated by `-->`. - The text that follows the timing information is the subtitle text itself, and it may span multiple lines. - Entries are separated by a blank line. ### How do VTT Formats Work? WEBVTT (Web Video Text Tracks), which is a standard format for displaying timed text tracks (such as subtitles or captions) within HTML5 video. The syntax is similar to SRT but has some differences: - The file should begin with the header WEBVTT. - Timing is done with a period (`.`) separating seconds and milliseconds instead of a comma (`,`). - No blank lines are needed between entries. - No index numbers are required. - This format is supported by many modern browsers and can be ussed with the HTML5 `` element to add subtitles to a `