Speech to Text and Audio Cleansing with AI Improving Transcription Accuracy |
Author(s): |
| Dipti Deepak Suryavanshi , PES Modern College of Engineering, Pune; Mr. Shripad S. Bhide, PES Modern College of Engineering, Pune |
Keywords: |
| Speech-to-Text (STT) Systems, AI-Powered Speech Recognition, Transcription Accuracy, Deep Learning in Speech Processing, Noise Reduction Techniques, Speaker Diarization, Adaptive Audio Filtering, Word Error Rate (WER), Signal-to-Noise Ratio (SNR), Voice Interaction Technology |
Abstract |
|
Speech-to-text technology has become a cornerstone in many areas such as virtual assistants, automatic transcription, accessibility tools, and real-time communication solutions. Our growing reliance on voice interactions has sparked significant advancements in AI-powered speech recognition systems. However, achieving high transcription accuracy presents a real challenge, especially with external factors like background noise, varying accents, crowded environments, and subpar audio quality. The rise of deep learning and AI-enhanced audio cleaning techniques—like noise reduction, echo cancellation, and speaker diarization—has notably improved transcription accuracy. Even so, we're still not quite reaching that human-like transcription quality, particularly in noisy settings or with low-resource languages. This paper takes a deep dive into AI-driven speech-to-text (STT) systems, exploring their evolution and impact on transcription accuracy. We compare various AI-driven STT models, such as OpenAI Whisper, Mozilla DeepSpeech, IBM Watson Speech-to-Text, and Google Speech-to-Text API, analyzing how they perform in different environments. By testing these models under various acoustic conditions, we assess their capabilities against the real-world problems they might face. We evaluate transcription quality using key metrics like Word Error Rate (WER), Signal-to-Noise Ratio (SNR), and latency issues. Additionally, we look into AI-based audio preprocessing methods to see how they help reduce errors and boost transcription quality. This includes examining deep learning noise reduction techniques, adaptive filtering, and speech enhancement models to understand how audio processing before transcription can mitigate external disturbances. Moreover, this study delves into the practical uses of AI-powered STT across various sectors, including healthcare for medical transcriptions, education for lecture notes, and customer support for automated call center services. |
Other Details |
|
Paper ID: IJSRDV13I30173 Published in: Volume : 13, Issue : 3 Publication Date: 01/06/2025 Page(s): 274-281 |
Article Preview |
|
|
|
|
