信息通信技术与政策

信息通信技术与政策

信息通信技术与政策 ›› 2026, Vol. 52 ›› Issue (9): 69-77.doi: 10.12267/j.issn.2096-5931.2026.09.010

技术与标准 上一篇    下一篇

面向实时语音转写场景的流式与非流式语音识别性能对比研究

Performance comparison of streaming and non-streaming speech recognition for real-time transcription scenarios

李鲲程1, 葛宏伟1, 郝航1, 范春梅2   

  1. 1 中国信息通信研究院信息管理中心, 北京 100191
    2 北京邮电大学信息与通信工程学院, 北京 100876
  • 收稿日期:2026-04-03 出版日期:2026-09-25 发布日期:2026-09-30
  • 通讯作者: 范春梅
  • 作者简介:
    李鲲程,中国信息通信研究院信息管理中心高级工程师,主要从事人工智能系统架构设计与优化、前沿技术落地验证、信息化建设项目管理和数据分析等方面的工作;
    葛宏伟,中国信息通信研究院信息管理中心项目主管,主要从事人工智能应用需求场景分析、应用系统设计与全流程管理等方面的工作;
    郝航,中国信息通信研究院信息管理中心工程师,主要从事计算平台的规划建设与运维等方面的工作

LI Kuncheng1, GE Hongwei1, HAO Hang1, FAN Chunmei2   

  1. 1 Information Management Center, China Academy of Information and Communications Technology, Beijing 100191, China
    2 School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China
  • Received:2026-04-03 Online:2026-09-25 Published:2026-09-30
  • Contact: FAN Chunmei

摘要:

为满足会议机器同传显示场景的需求,实时语音转写需满足低时延、高准确率、强可读性的要求。为此,通过构建系统化的评估指标体系和离线流式模拟评测方法,对比流式(Sherpa-onnx模型)与非流式(Faster-Whisper模型)两种语音识别方案的性能差异,为相关领域研究提供参考。试验结果表明,流式语音识别方案配合轻量级语义后处理策略,在整体性能表现上更适应工业级会议系统对实时语音转写的综合要求。

关键词: 实时语音转写, 滑动窗口, 语义后处理

Abstract:

To meet the requirements of real-time speech transcription display in conference simultaneous interpretation scenarios, the system must achieve low latency, high accuracy, and strong readability. For this purpose, systematic evaluation metrics and an offline streaming simulation assessment method are established to compare the performance differences between streaming (Sherpa-onnx) and non-streaming (Faster-Whisper) speech recognition solutions, offering a reference for research in related fields. Experimental results indicate that the streaming speech recognition approach, when coupled with a lightweight semantic post-processing strategy, exhibits overall performance that better satisfies the comprehensive demands of real-time speech transcription in industrial-grade conference systems.

Key words: real-time speech transcription, sliding window, semantic post-processing

中图分类号: