Back to Skills

scribe

Reference skill for Zoom AI Services Scribe. Use after routing to a transcription workflow when handling uploaded or stored media, Build-platform JWT auth, fast mode transcription, batch jobs, or transcript pipeline design.

23,453stars2,830forksUpdated 8/13/2026

Security Assessment

Safe(93/100)
Security Score93/100

About scribe

This reference skill covers Zoom AI Services Scribe, a transcription API for turning uploaded or stored media into text. It helps developers design transcript pipelines around synchronous single-file (fast mode) transcription, asynchronous batch jobs, webhook-driven status updates, and Build-platform JWT authentication, and it clarifies where Scribe fits versus real-time media streaming.

The skill documents the core workflow: generate an HS256 Build-platform JWT, choose fast mode for one short file or batch mode for stored archives and large sets, submit the request, poll job/file status or receive webhooks for batch jobs, then persist and post-process the transcript JSON. It maps the endpoint surface (transcribe, jobs create/list/inspect/cancel, per-file results) and includes practical operational guidance: real deployed-sample timing observations, a note that a frontend 504 with a backend 200 is a browser/edge timeout race rather than a transcription failure, a recommendation to wrap fast mode in async request/polling for hosted UIs, and a browser-microphone pseudo-streaming pattern (short chunked uploads) explicitly framed as a fallback demo rather than a substitute for real-time streaming. It also documents credential-naming drift across Zoom docs and advises verifying labels in the portal. Its auth guidance is defensive — keep JWTs short-lived (one hour or less) and sign server-side.

It targets backend and application developers building on-demand clip transcription, batch transcription of stored archives, webhook-driven ETL pipelines, and compliance/QA transcript workflows. The content is documentation and code-pattern reference grounded in the official Zoom AI Services docs and the scribe-quickstart sample.

FAQ

What can Scribe transcribe?

Uploaded or stored media files, either synchronously in fast mode for one short file or asynchronously via batch jobs for stored archives and large sets. It is not a documented real-time streaming API.

How is authentication handled?

Via a Build-platform HS256 JWT bearer token. The skill recommends keeping expiration to one hour or less and signing server-side, and notes Zoom docs use inconsistent credential labels (API key/secret, SDK key/secret, Build-platform credentials) for the same issuer/secret pair.

What are the fast-mode limits?

The formal fast-mode API limits are 100 MB and 2 hours, but hosted browser flows can still hit edge timeouts; the skill advises an async request/polling wrapper and preferring batch mode for larger or less predictable media.

Can I do live microphone transcription?

Only via a pseudo-streaming fallback of short chunked uploads (roughly 5-10 second chunks), which the skill frames as a demo pattern. For true live ingestion it routes you to the RTMS skill.

What are typical use cases?

On-demand clip transcription, batch transcription of stored archives, webhook-driven ETL that writes transcripts to a database or search index, and offline compliance or QA workflows needing timestamps and speaker hints.

All Files

11 files
SKILL.md5.8 KB
View
concepts/auth-and-processing-modes.md3.2 KB
View
references/api-reference.md2.6 KB
View
references/versioning-and-drift.md2.0 KB
View
scenarios/high-level-scenarios.md3.1 KB
View
examples/fast-mode-node.md1.8 KB
View
references/samples-validation.md2.0 KB
View
RUNBOOK.md4.2 KB
View
troubleshooting/common-drift-and-breaks.md4.6 KB
View
examples/batch-webhook-pipeline.md1.6 KB
View
references/environment-variables.md1.5 KB
View

Install scribe

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill