Information and Communications Technology and Policy

Information and Communications Technology and Policy

Information and Communications Technology and Policy ›› 2026, Vol. 52 ›› Issue (9): 69-77.doi: 10.12267/j.issn.2096-5931.2026.09.010

Previous Articles     Next Articles

Performance comparison of streaming and non-streaming speech recognition for real-time transcription scenarios

LI Kuncheng1, GE Hongwei1, HAO Hang1, FAN Chunmei2   

  1. 1 Information Management Center, China Academy of Information and Communications Technology, Beijing 100191, China
    2 School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China
  • Received:2026-04-03 Online:2026-09-25 Published:2026-09-30
  • Contact: FAN Chunmei

Abstract:

To meet the requirements of real-time speech transcription display in conference simultaneous interpretation scenarios, the system must achieve low latency, high accuracy, and strong readability. For this purpose, systematic evaluation metrics and an offline streaming simulation assessment method are established to compare the performance differences between streaming (Sherpa-onnx) and non-streaming (Faster-Whisper) speech recognition solutions, offering a reference for research in related fields. Experimental results indicate that the streaming speech recognition approach, when coupled with a lightweight semantic post-processing strategy, exhibits overall performance that better satisfies the comprehensive demands of real-time speech transcription in industrial-grade conference systems.

Key words: real-time speech transcription, sliding window, semantic post-processing

CLC Number: