Skip to main content

The Babel Group

Share

By: Dr. Rachel Khasky-Levy, SLPD, CCC-SLP

In short: standard speech-to-text tools like Google Read&Write run on a general recognition engine trained on typical speech, so word error rates can run very high for students with dysarthria, apraxia, or other atypical speech patterns. Voiceitt trains on the student’s own voice instead, which can meaningfully lower that error rate for students whose standard tools weren’t built for. This article shows a simple way to measure that difference, called word error rate (WER), and turns it into two kinds of IEP goals: one any Voiceitt candidate can work toward, and one reserved for students who can also review and correct their own output.

When a student uses speech recognition technology, the numbers matter. Not because a perfect transcript is the goal, but because the right tool can turn a communication breakdown into a communication opportunity. This picks up once Voiceitt has become a familiar listener for the student through the training process (see the Taking data blog).

Why Google Read&Write doesn’t work for every candidate

Google Read&Write is a widely used Chrome extension with a Talk&Type feature that converts speech to text inside Google Docs and Slides. It’s a solid tool for many students, and it’s worth understanding how it works before explaining where it falls short.

Talk&Type runs on Google’s general speech recognition engine, the same one behind standard Chromebook dictation. That engine is trained on typical speech patterns from a broad population. It listens for clear articulation, standard pacing, and consistent pronunciation.

That’s exactly where the gap shows up for students with dysarthria, apraxia, or other atypical speech patterns. The recognition engine wasn’t trained on their speech, so it doesn’t adapt to it. Every session starts from the same generic model. Word error rate for these students varies widely depending on the severity and type of speech pattern, and for some students with more severe patterns it can be very high. That’s not a tool problem exactly; it’s a training data problem. The engine simply has no reference point for how this student talks.

Voiceitt takes a different approach. Instead of relying on a generic model, it trains on the individual student’s own voice, building a personalized recognition profile over time. That’s why it can bring word error rate down for students the standard tools weren’t built for, though the size of that improvement will differ from student to student.

A simple way to compare two speech recognition tools

You don’t need a research background to measure this. Here’s a simplified process any team can run.

  1. Pick a set of sentences or short passages the student will say (aim for at least 20 to 50 words total for a meaningful sample).
  2. Have the student use standard speech-to-text (like Read&Write’s Talk&Type) to produce the text. Record the output exactly as generated.
  3. Have the student use Voiceitt to produce the text for the same passage.
  4. Compare each output word by word against what the student actually intended to say.
  5. Count the errors: any word that’s wrong, missing, or added that shouldn’t be there.

Calculating word error rate (WER)

The formula looks technical but the concept is simple: what percentage of words came out wrong.

WER = (number of word errors ÷ total number of words intended) × 100

For example, if a student intended to say a 20-word sentence and 18 words came out wrong using standard speech-to-text, that’s a 90 percent WER for that sample. If the same sentence produced 10 errors using Voiceitt, that’s a 50 percent WER. These numbers will vary by student; this is just to illustrate the math.

To find the percent reduction between the two tools, use this formula:

Percent reduction = ((baseline WER − new WER) ÷ baseline WER) × 100

Using the numbers above as an illustration only: (90 − 50) ÷ 90 × 100 = about 44 percent reduction. Each student’s actual baseline and reduction will depend on their individual speech pattern and needs to be measured, not assumed.

Building IEP goals around this data

The most important thing to get right when writing these goals: don’t build in an assumption you haven’t tested. A student’s ability to recognize and correct transcription errors is a separate skill from their ability to produce clearer speech output with Voiceitt. Mixing the two into one goal risks writing something the student can’t achieve for reasons that have nothing to do with the technology.

Here’s how to split it.

Core goal, achievable for any Voiceitt candidate:

Student will use Voiceitt for academic verbal communication and/or dictation, achieving a word error rate reduced by at least 30 percent compared to standard speech recognition software, as measured by comparative baseline data collected across 4 consecutive sessions.

This goal reflects what the technology does on its own. It doesn’t depend on the student’s editing skills, reading level, or independence with correction strategies. It’s measurable, it’s realistic, and it’s true to what Voiceitt is actually built to do.

Self-correction goal, only for students with assessed capacity to review and edit their own output:

Given Voiceitt-dictated text containing transcription errors, the student will independently identify and correct errors to produce a final written product at grade-level accuracy, in 4 out of 5 writing assignments.

This goal layers on top of the first one, but it should only go into an IEP once an SLP or the team has confirmed the student can actually review and edit text. Writing this goal without that confirmation sets up a target the student may not be positioned to hit, and that’s a goal quality issue, not a technology issue.

The bottom line

Voiceitt isn’t a perfect transcription tool, and it doesn’t need to be to make a real difference. For students for whom standard speech recognition produces a wall of errors, cutting that error rate by even a third can be the difference between a student who gives up on verbal participation and one who engages consistently in class discussion and written work. The goal isn’t flawless output. It’s a measurable, documented reduction in the communication breakdown that’s been getting in the way.

To obtain a quote or learn more about different funding pathways and our pricing model, please reach out to support@thebabelgroup.com


Share

Leave a Reply

Your email address will not be published. Required fields are marked *