High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Speech to Text and Audio Cleansing with AI Improving Transcription Accuracy

Author(s):

Dipti Deepak Suryavanshi , PES Modern College of Engineering, Pune; Mr. Shripad S. Bhide, PES Modern College of Engineering, Pune

Keywords:

Speech-to-Text (STT) Systems, AI-Powered Speech Recognition, Transcription Accuracy, Deep Learning in Speech Processing, Noise Reduction Techniques, Speaker Diarization, Adaptive Audio Filtering, Word Error Rate (WER), Signal-to-Noise Ratio (SNR), Voice Interaction Technology

Abstract

Speech-to-text technology has become a cornerstone in many areas such as virtual assistants, automatic transcription, accessibility tools, and real-time communication solutions. Our growing reliance on voice interactions has sparked significant advancements in AI-powered speech recognition systems. However, achieving high transcription accuracy presents a real challenge, especially with external factors like background noise, varying accents, crowded environments, and subpar audio quality. The rise of deep learning and AI-enhanced audio cleaning techniques—like noise reduction, echo cancellation, and speaker diarization—has notably improved transcription accuracy. Even so, we're still not quite reaching that human-like transcription quality, particularly in noisy settings or with low-resource languages. This paper takes a deep dive into AI-driven speech-to-text (STT) systems, exploring their evolution and impact on transcription accuracy. We compare various AI-driven STT models, such as OpenAI Whisper, Mozilla DeepSpeech, IBM Watson Speech-to-Text, and Google Speech-to-Text API, analyzing how they perform in different environments. By testing these models under various acoustic conditions, we assess their capabilities against the real-world problems they might face. We evaluate transcription quality using key metrics like Word Error Rate (WER), Signal-to-Noise Ratio (SNR), and latency issues. Additionally, we look into AI-based audio preprocessing methods to see how they help reduce errors and boost transcription quality. This includes examining deep learning noise reduction techniques, adaptive filtering, and speech enhancement models to understand how audio processing before transcription can mitigate external disturbances. Moreover, this study delves into the practical uses of AI-powered STT across various sectors, including healthcare for medical transcriptions, education for lecture notes, and customer support for automated call center services.

Other Details

Paper ID: IJSRDV13I30173
Published in: Volume : 13, Issue : 3
Publication Date: 01/06/2025
Page(s): 274-281

Article Preview

Download Article