# Using google's voice recognition to convert audio files to text

**URL:** <https://boards.straightdope.com/t/using-googles-voice-recognition-to-convert-audio-files-to-text/706926>\
**Category:** Factual Questions\
**Created:** [December 11, 2014, 7:55am UTC](https://boards.straightdope.com/t/using-googles-voice-recognition-to-convert-audio-files-to-text/706926 "2014-12-11T07:55:58Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![The\_Sheikh](https://avatars.discourse-cdn.com/v4/letter/t/c67d28/32.png) [@The\_Sheikh](https://boards.straightdope.com/u/The_Sheikh)\
**Post date:** [December 11, 2014, 7:55am UTC](https://boards.straightdope.com/t/using-googles-voice-recognition-to-convert-audio-files-to-text/706926/1 "2014-12-11T07:55:58Z")

</div>

Google’s api for voice recognition is more accurate than most voice recognition software that I have tried, including [Nuance’s Dragon](http://www.nuance.com/dragon/index.htm). In addition, it recognizes a multitude of languages.

Is there any way I can exploit Google’s API to convert audio files into text? I want to upload lengthy audio files which would be then converted to text; not speak into a microphone.

---

<div class="post-metadata">

**Author:** ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)\
**Post date:** [December 11, 2014, 8:43am UTC](https://boards.straightdope.com/t/using-googles-voice-recognition-to-convert-audio-files-to-text/706926/2 "2014-12-11T08:43:48Z")

</div>

You could try YouTube’s automatic captions feature.

---

<div class="post-metadata">

**Author:** ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)\
**Post date:** [December 11, 2014, 8:53am UTC](https://boards.straightdope.com/t/using-googles-voice-recognition-to-convert-audio-files-to-text/706926/3 "2014-12-11T08:53:06Z")

</div>

Or you can try one of the following API demos:

> **[Chrome Browser](https://www.google.com/intl/en/chrome/demos/speech.html)**
>
> Google Chrome is a browser that combines a minimal design with sophisticated technology to make the web faster, safer, and easier.

> **[SpeechRecognition](https://pypi.org/project/SpeechRecognition/)**
>
> Library for performing speech recognition, with support for several engines and APIs, online and offline.

> <https://gist.github.com/alotaiba/1730160>

But even if the underlying engine is great (and it is), you don’t get the full UI of something like Nuance. You can’t easily go back and correct mistakes, customize how it’s trained, etc.
