-70%

Median Processing Turnaround

12,000

Concurrent Transcription Jobs Supported

Portrait of Yuna Park, CTO and Co-Founder at Driftcast

Yuna Park,

Yuna Park

CTO & Co-Founder

Geometric prism visualization demonstrating parallel inference fan-out for high-throughput asynchronous media processing.

The Problem: Transcription Was a Sequential Bottleneck in the Upload Pipeline

Driftcast’s async video platform runs every uploaded recording through a transcription and auto-chapter pipeline before the video is marked ready for sharing. The legacy pipeline processed audio chunks sequentially through a single Whisper-class model instance per upload, meaning a 30-minute recording took roughly 9 minutes of inference time before chaptering and search-indexing could even begin.

The core problem was that chunk-level parallelism was never implemented; the pipeline treated each upload as a single monolithic inference job rather than a set of independently processable segments. At peak upload volume, the processing queue backed up by over 40 minutes, directly delaying the moment a recorded video became shareable.

The Implementation: Meridian Parallel Inference Fan-Out

Driftcast re-architected the pipeline around Meridian’s fan-out inference primitive, which splits a single upload into overlapping 30-second audio segments, dispatches them as a batch of concurrent inference requests, and reassembles ordered transcript output using segment timestamps and a small overlap-deduplication pass to handle word boundaries split across chunks.

The hard part was never running Whisper faster on a single chunk. It was reliably fanning out hundreds of chunks per minute and reassembling them in the right order without duplicated or dropped words at the seams.

Meridian’s autoscaling worker pool absorbs upload bursts by scaling transcription workers independently from the rest of Driftcast’s media-processing stack, which previously scaled as a single coupled service and over-provisioned compute for stages that did not need it.

Results: Measured Infrastructure Impact

Median processing turnaround for a 30-minute recording fell from approximately 9 minutes to under 3 minutes, a 70% reduction. The platform now sustains 12,000 concurrent transcription jobs during peak periods without queue backup, compared to a prior ceiling of roughly 1,800 concurrent jobs before queueing delays became customer-visible.

Ready to route your first payload?

Get your first API key and start routing production traffic today.

Aquire $129

Aquire $129

Aquire $129

Create a free website with Framer, the website builder loved by startups, designers and agencies.