IBM Says Granite Speech 5.0 Transcribes 3.5 Hours of Speech in One Second
SMRTR summary
IBM released two compact English speech recognition models on August 25, 2026, claiming a transcription speed no open model has matched before. The Granite Speech 5.0 TurboCTC models each have 470 million parameters and can process more than 3.5 hours of audio per second on a single NVIDIA H200 GPU, running over 12,600 times faster than real time, making fast, affordable transcription far more accessible.
SMRTR provides this summary for quick context. The original article belongs to Unite AI.
Read the original article