# Voice to Text for Meeting Minutes

**URL:** <https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226>\
**Category:** In My Humble Opinion\
**Created:** [July 17, 2019, 8:34pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226 "2019-07-17T20:34:46Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bone](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/bone/32/407_2.png) [@Bone](https://boards.straightdope.com/u/Bone)\
**Post date:** [July 17, 2019, 8:34pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/1 "2019-07-17T20:34:46Z")

</div>

I’m looking for a solution that would allow the capture of meeting minutes without having for a meeting Secretary to record, and manually type out meeting minutes. Is there any system that folks are aware of that would do this? Think board meetings with large numbers of participants.

I have used an audio recording and then played it through my computer and had google docs use voice to text and have it transcribe, however this doesn’t capture who is actually talking, or do punctuation very well.

I’m willing to pay if there is a system out there that would do it. Ideally all the speakers have designated microphones and based on whose mic is being used it would attribute that text to that person.

---

<div class="post-metadata">

**Author:** ![control-z](https://avatars.discourse-cdn.com/v4/letter/c/eada6e/32.png) [@control-z](https://boards.straightdope.com/u/control-z)\
**Post date:** [July 17, 2019, 8:39pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/2 "2019-07-17T20:39:37Z")

</div>

From everything I’ve seen, no such technology exists. I’ve used Google voice recognition a lot, and it’s only mediocre accuracy at best. Some of the typos are hilariously wrong, others are subtle. And that’s just my voice, after years of trying. I don’t think there is any way voice recognition is going to accurately recognize what different speakers are saying.

---

<div class="post-metadata">

**Author:** ![Bone](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/bone/32/407_2.png) [@Bone](https://boards.straightdope.com/u/Bone)\
**Post date:** [July 18, 2019, 8:44pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/3 "2019-07-18T20:44:28Z")

</div>

I’ve been looking at a product called Voicea, but can’t find a lot of user info. Anyone?

---

<div class="post-metadata">

**Author:** ![Dewey\_Finn](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dewey_finn/32/4222_2.png) [@Dewey\_Finn](https://boards.straightdope.com/u/Dewey_Finn)\
**Post date:** [July 18, 2019, 9:50pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/4 "2019-07-18T21:50:10Z")

</div>

Take a look at [Otter](https://otter.ai/) and [Trint](https://trint.com/).which are designed for transcription, either real time or in a short time. Both were mentioned in recent New York Times articles in which Times journalists described the technology they use. I really doubt either will identify who is speaking, though.

---

<div class="post-metadata">

**Author:** ![Pork\_Rind](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pork_rind/32/111_2.png) [@Pork\_Rind](https://boards.straightdope.com/u/Pork_Rind)\
**Post date:** [July 18, 2019, 9:54pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/5 "2019-07-18T21:54:49Z")

</div>

I’ve used Trint a couple of times to capture the conversation from working sessions that I was leading where I couldn’t bring along a note taker. I thought the quality was good, although there were difficult to follow sections where several people were talking at once. I don’t recall it had any way to identify the speaker. I had to go back later and do that by ear.

---

<div class="post-metadata">

**Author:** ![Tee](https://avatars.discourse-cdn.com/v4/letter/t/73ab20/32.png) [@Tee](https://boards.straightdope.com/u/Tee)\
**Post date:** [July 19, 2019, 12:00am UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/6 "2019-07-19T00:00:39Z")

</div>

If recorded, they wouldn’t be minutes, they’d be transcripts. Legal minutes record board actions and not every spoken word.

---

<div class="post-metadata">

**Author:** ![Bone](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/bone/32/407_2.png) [@Bone](https://boards.straightdope.com/u/Bone)\
**Post date:** [July 19, 2019, 7:45am UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/7 "2019-07-19T07:45:20Z")

</div>

The purpose is to assist in minute taking. Some minutes can be very short, others much more verbose. To reduce manual typing a transcription can be a good starting point.

---

<div class="post-metadata">

**Author:** ![Tee](https://avatars.discourse-cdn.com/v4/letter/t/73ab20/32.png) [@Tee](https://boards.straightdope.com/u/Tee)\
**Post date:** [July 19, 2019, 2:30pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/8 "2019-07-19T14:30:10Z")

</div>

I don’t blame you for trying to make the process more efficient. When I do it, the job is to condense and summarize verbosity (for hours) and use it rather sparingly in describing board actions. This creates a formal document that satisfies legal obligations. A transcript would be a separate documentation that is not legally required, but which now legally exists, and not everyone wants that. Just a heads up.

---

<div class="post-metadata">

**Author:** ![Bone](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/bone/32/407_2.png) [@Bone](https://boards.straightdope.com/u/Bone)\
**Post date:** [July 19, 2019, 10:01pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/9 "2019-07-19T22:01:18Z")

</div>

I was looking at Amazon Transcribe, but then I exceed my technical knowledge. It looks like it’s a service that is reasonably priced, but I don’t know how to actually test it. It talks about development in AWS, and APIs but I don’t know how to interpret it.

After experimenting with Voicea, I don’t think that would work. I also tried Trint and it seems to work well so that’s good.

---

<div class="post-metadata">

**Author:** ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)\
**Post date:** [July 20, 2019, 12:42am UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/10 "2019-07-20T00:42:05Z")

</div>

Came in here to suggest Trint. Play with it some; it’s actually very good at what it does. Sure, you’ll have to do a little bit of manual cleanup afterward, but a lot less than typing it all out by hand.

And they do have speaker separation as a feature, but last I tried, it was less than stellar:

[https://support.trint.com/hc/en-us/articles/360000235517-Speaker-Separation](https://support.trint.com/hc/en-us/articles/360000235517-Speaker-Separation)

---

<div class="post-metadata">

**Author:** ![Madam\_Librarian](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/madam_librarian/32/3123_2.png) [@Madam\_Librarian](https://boards.straightdope.com/u/Madam_Librarian)\
**Post date:** [July 21, 2019, 12:46am UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/11 "2019-07-21T00:46:08Z")

</div>

A good number of voice-to-text programs require one to _teach_ the software to interpret voices, so you’ll have to consider the set-up time involved (this includes all of the speakers providing speaking samples) for the sake of accuracy. Moreover, because it’s often difficult for them to distinguish among multiple voices, there could be several hours of editing the transcript, identifying who said what. And, it’s miserable when people talk over each other. Most, like [Dragon](https://www.nuance.com/dragon.html), and [Transcribe](https://transcribe.wreally.com/) are going to present the same types of problems you have with Google Docs.

Working in a university oral history program for decades, this was an ongoing discussion that, upon several attempts to use this type of software, always resulted in our decision to return to more traditional transcription from the sound recordings.

---

<div class="post-metadata">

**Author:** ![don\_t\_ask](https://avatars.discourse-cdn.com/v4/letter/d/e68b1a/32.png) [@don\_t\_ask](https://boards.straightdope.com/u/don_t_ask)\
**Post date:** [July 21, 2019, 1:54am UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/12 "2019-07-21T01:54:46Z")

</div>

Arbie’s dragon fought a porpoise and bin elbow to achieve excellent results - as toucan see here.

---

<div class="post-metadata">

**Author:** ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)\
**Post date:** [July 24, 2019, 4:31pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/13 "2019-07-24T16:31:23Z")

</div>

> [@Madam\_Librarian](#):
>
> A good number of voice-to-text programs require one to _teach_ the software to interpret voices, so you’ll have to consider the set-up time involved (this includes all of the speakers providing speaking samples) for the sake of accuracy. Moreover, because it’s often difficult for them to distinguish among multiple voices, there could be several hours of editing the transcript, identifying who said what. And, it’s miserable when people talk over each other. Most, like [Dragon](https://www.nuance.com/dragon.html), and [Transcribe](https://transcribe.wreally.com/) are going to present the same types of problems you have with Google Docs.
> 
> Working in a university oral history program for decades, this was an ongoing discussion that, upon several attempts to use this type of software, always resulted in our decision to return to more traditional transcription from the sound recordings.

Have you tried Trint? I’d be curious as to your thoughts on it, as someone who’s used similar stuff for years.

Machine learning has drastically improved speech recognition in the last 5-10 years, using new technology entirely different from the old Dragons and such. Even [speaker separation is making huge strides](https://www.youtube.com/watch?v=vW51cG1Ox98). This has to do with huge companies like Google and Amazon investing big-time in the tech, powering things like Alexa, Siri, and the Google Assistant, using machine-driven statistical analyses over tens of thousands of hours of recordings (an approach that wasn’t yet quite feasible in decades past).

For example, now YouTube can automatically caption (to maybe 80% accuracy?) uploaded videos in several languages with no prior training from the speaker(s), entirely for free. It’s not as user-friendly as Trint, but is still a very affordable way of getting semi-usable transcripts that you can edit in much less time than manually transcribing from scratch.

---

<div class="post-metadata">

**Author:** ![Kropotkin](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/kropotkin/32/18045_2.png) [@Kropotkin](https://boards.straightdope.com/u/Kropotkin)\
**Post date:** [July 25, 2019, 10:24pm UTC](https://boards.straightdope.com/t/voice-to-text-for-meeting-minutes/837226/14 "2019-07-25T22:24:21Z")

</div>

> [@Tee](#):
>
> If recorded, they wouldn’t be minutes, they’d be transcripts. Legal minutes record board actions and not every spoken word.

This is an important distinction and it’s not clear from the OP that it has been grasped. Minutes do not, should not, report who said what on what issue but motions and votes and directives to officers and the like. They don’t include committee reports, speeches, questions, and the random lunacy of meetings. A 3 hour strata council meeting might generate 2 pages of minutes, for example
