| [1] |
Radford A, Kim J W, Xu T, et al. Robust speech recognition via large-scale weak supervision[PP/OL]. V1. arXiv (2022-12-06) [2026-04-03]. https://doi.org/10.48550/arXiv.2212.04356.
|
| [2] |
Macháček D, Dabre R, Bojar O. Turning whisper into real-time transcription system[PP/OL]. V1. arXiv (2023-07-27) [2026-04-03]. https://doi.org/10.48550/arXiv.2307.14743.
|
| [3] |
Wang H Y, Hu G Q, Lin G D, et al. Simul-whisper: attention-guided streaming whisper with truncation detection[PP/OL]. V1. arXiv ( 2024-06-24) [2026-04-03]. https://doi.org/10.21437/Interspeech.2024-1814.
|
| [4] |
Tsunoo E, Kashiwagi Y, Watanabe S. Streaming transformer asr with blockwise synchronous beam search[C]// 2021 IEEE Spoken Language Technology Workshop (SLT). Shenzhen: IEEE Press, 2021: 22-29.
|
| [5] |
Narayanan A, Prabhavalkar R, Chiu C C, et al. Recognizing long-form speech using streaming end-to-end models[C]// 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). Singapore: IEEE Press, 2019: 920-927.
|
| [6] |
Wu C Y, Wang Y Q, Shi Y Y, et al. Streaming transformer-based acoustic models using self-attention with augmented memory[PP/OL]. V1. arXiv (2020-05-16) [2026-04-03]. https://doi.org/10.48550/arXiv.2005.08042.
|
| [7] |
Yao Z W, Guo L Y, Yang X Y, et al. Zipformer: a faster and better encoder for automatic speech recognition[PP/OL]. V1. arXiv (2023-10-17) [2026-04-03]. https://doi.org/10.48550/arXiv.2310.11230.
|
| [8] |
Chen G G, Chai S Z, Wang G B, et al. Gigaspeech:an evolving, multi-domain asr corpus with 10,000 hours of transcribed audio[PP/OL]. V1. arXiv (2021-06-13) [2026-04-03]. https://doi.org/10.21437/Interspeech.2021-1965.
|
| [9] |
Guhr O, Schumann A-K, Bahrmann F, et al. Fullstop: multilingual deep models for punctuation prediction[C]//Proceedings of the Swiss Text Analytics Conference 2021. Winterthur: CEUR Workshop Proceedings, 2021.
|
| [10] |
Panayotov V, Chen G, Povey D, et al. Librispeech: an asr corpus based on public domain audio books[C]// 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). South Brisbane: IEEE Press, 2015: 5206-5210.
|
| [11] |
Hernandez F, Nguyen V, Ghannay S, et al. TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptation[C]// International Conference on Speech and Computer. Cham: Springer International Publishing, 2018: 198-208.
|
| [12] |
Carletta J. Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus[J]. Language Resources and Evaluation, 2007, 41 (2): 181-190.
doi: 10.1007/s10579-007-9040-x
URL
|
| [13] |
Li J Y. Recent advances in end-to-end automatic speech recognition[PP/OL]. V1. arXiv (2021-11-02) [2026-04-03]. https://doi.org/10.48550/arXiv.2111.01690.
|
| [14] |
Moritz N, Hori T, Le J. Streaming automatic speech recognition with the transformer model[C]// ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Barcelona: IEEE Press, 2020: 6074-6078.
|
| [15] |
Miller R B. Response time in man-computer conversational transactions[C]// Proceedings of the December 9-11, 1968, Fall Joint Computer Conference, Part I. San Francisco: Thompson Book Company, 1968: 267-277.
|
| [16] |
Graves A. Sequence transduction with recurrent neural networks[PP/OL]. V1. arXiv (2012-11-14) [2026-04-03]. https://doi.org/10.48550/arXiv.1211.3711.
|
| [17] |
Morris A C, Maier V, Green P D. From WER and RIL to MER and WIL: improved evaluation measures for connected speech recognition[C]//Interspeech 2004. Jeju: ISCA, 2004: 2765-2768.
|
| [18] |
李鲲程, 郝航, 温仰飞, 等. 基于首Token和平均Token响应时间的大模型负载测试方法研究[J]. 通信管理与技术, 2025 (5): 6-10.
|
| [19] |
Crankshaw D, Wang X, Zhou G, et al. Clipper: A {Low-Latency} online prediction serving system[C]// 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). Boston: USENIX Association, 2017: 613-627.
|