Speech recognition in aphasia therapy: what it can and cannot do

It closes the loop on a spoken attempt when no clinician is in the room. It also fails in specific, predictable ways that are worth knowing before you rely on it.

By Frederick F. Weiner, Ph.D., speech-language pathologist. Former Professor in Charge of Speech Pathology, The Pennsylvania State University. Last reviewed 15 August 2026.

The problem it solves

In a therapy room, a clinician hears every attempt and responds. At home, most practice software cannot hear anything. A person is shown a picture, says a word aloud, and then taps a button to reveal whether they were right. Nothing checked what they actually said.

Speech recognition changes that. The program listens, decides whether the target word was produced, and responds immediately. That matters because speaking is the skill being trained. Tapping the correct picture is a different task, and practising it does not practise speech.

Where it works well

Speech recognition is most useful when someone knows roughly what they want to say and struggles to retrieve or produce the word, and their speech is still reasonably intelligible. In that situation it supplies two things that are hard to get at home: an immediate answer, and the ability to practise many attempts without another person present.

The strongest published demonstration of the approach is not ours. In a phase II trial reported in eClinicalMedicine in 2024, University College London researchers tested iTalkBetter, a gamified app built around a speech recogniser that verifies naming attempts in real time. Twenty-seven people with post-stroke aphasia practised for around 90 minutes a day over six weeks and improved naming of 200 commonly used words by about 13 percent. That is good evidence that real-time recognition feedback can drive measurable gains.

It is worth being clear about what that does and does not mean for anyone reading about it, because the 2024 press coverage was substantial and people do go looking. iTalkBetter is a research app, not a product you can buy. As of August 2026 University College London describes it as “currently under development,” and the italkbetter.com domain no longer resolves to the project. The trial was also conducted in the United Kingdom, and we have found no published information on how the recogniser performs with American English or with other accents — which matters, given that recognition accuracy is not uniform across speakers.

Where it fails

Speech recognisers are built mainly on typical speech. Conditions that change how speech sounds degrade their accuracy, and in rehabilitation those conditions are common rather than rare.

  • Severe aphasia. When productions are highly variable or consist largely of non-words, there is often nothing for a recogniser to match, and constant rejection is discouraging rather than instructive.
  • Apraxia of speech. Attempts are inconsistent by definition. The same word may come out differently on consecutive tries, which is exactly the pattern recognisers handle worst. Research in this area is framed as feasibility work rather than settled practice.
  • Dysarthria. Weakness or poor coordination of the speech muscles changes timing and articulation across the board. Accuracy drops accordingly.
  • Background noise and microphone position. The most common cause of a correct answer being marked wrong is environmental, not neurological. A television in the room or a microphone across the desk will produce errors in anyone.
  • Accents and non-native English. Recognition accuracy is not uniform across speakers, and people whose speech differs from the training data get worse results through no fault of their own.

There is also a subtler risk. A recogniser can accept a word that was produced poorly but recognisably, and reject one that was produced well but quietly. Neither judgement is clinical. Speech recognition is feedback during practice — it is not an assessment, and it should not be read as one.

Getting better results from it

  • Use a headset microphone rather than the one built into a laptop or tablet. This single change fixes most complaints.
  • Practise in a quiet room with the television and radio off.
  • Wait for the program to start listening before speaking. Beginning too early clips the first sound of the word, which is often what the recogniser needs most.
  • Do not argue with it. If an attempt was clearly correct and was rejected, move on. Repeating an attempt five times to satisfy a microphone is frustrating and is not therapy.
  • Tell the clinician what it keeps rejecting. A pattern of consistent rejections on particular sounds is genuinely useful information for a speech-language pathologist — which is a better use of it than a score.

How Parrot Software uses it

Our speaking programs listen through a microphone and give instant feedback on the spoken response. Speech recognition runs across different kinds of speaking task rather than picture naming alone, because the same subscription covers speech, reading, memory, vocabulary, function, and cognition — and different people need different things.

In the study of our software by Corwin and colleagues at Texas Tech University Health Sciences Center, six people with chronic aphasia used eight of these programs in a 32-hour protocol and showed statistically significant improvement in naming words that were never trained. Six participants and no control group is a small study, and it did not isolate speech recognition as the active ingredient. The full text and its limitations are on our research page.

Programs that do not involve speaking do not use the microphone at all. If the goal is reading comprehension or memory, speech recognition is not the relevant feature and we would not claim otherwise.

Who should probably not rely on it

If speech is currently very limited, highly variable, or difficult for a familiar listener to understand, speech-recognition feedback will frustrate more than it helps, and a different approach is a better use of the effort. That is worth establishing during a free trial rather than after paying for a subscription. It is also exactly the kind of question a speech-language pathologist can answer in one session.

Common questions

What does speech recognition actually do in aphasia therapy software?
It listens to a spoken response through a microphone and decides whether the intended word was produced, then gives immediate feedback. The point is not to grade anyone. It is to close the loop on a practice attempt straight away, so a person practising alone at home gets a response rather than silence.
Does speech recognition work for everyone with aphasia?
No. It works best for people whose speech is reasonably intelligible and whose main difficulty is retrieving words. It is much less reliable for severe aphasia, for apraxia of speech, and for dysarthria, because those conditions change the sound of speech in ways general-purpose recognisers handle poorly. If speech is very variable or hard for a familiar listener to understand, expect frequent misrecognition.
Why does it sometimes mark a correct answer wrong?
Common causes are background noise, a microphone too far from the mouth, speaking before the program is listening, and speech patterns that differ from those the recogniser was trained on. Regional accents and non-native English can also reduce accuracy. A headset microphone in a quiet room removes most of these problems.
Should I trust speech recognition to tell me whether speech is improving?
Treat it as practice feedback, not as measurement. A recogniser saying yes or no on individual attempts is useful during practice. It is not an assessment of speech, and it cannot tell you what to work on next. That judgement belongs to a speech-language pathologist.
Can I download iTalkBetter?
Not currently. iTalkBetter is a research app developed at University College London, and as of August 2026 UCL describes it as under development. Its former website domain no longer resolves to the project. The 2024 trial results are real and encouraging, but the app is not available to buy or download, and its trial was conducted in the United Kingdom with no published information about performance on American English.
Is speech recognition better than tapping or matching exercises?
For practising speech production, yes, because saying a word aloud is the skill being trained and tapping a picture is not. For other goals such as reading comprehension, memory, or cognitive tasks, speech recognition is beside the point. The right question is what the practice is meant to achieve.

References

  • Efficacy of a gamified digital therapy for speech production in people with chronic aphasia (iTalkBetter): behavioural and imaging outcomes of a phase II item-randomised clinical trial. eClinicalMedicine. 2024. PubMed 38685927
  • Feasibility of Automatic Speech Recognition for Providing Feedback During Tablet-Based Treatment for Apraxia of Speech Plus Aphasia. PubMed 31306595
  • Corwin, M., Wells, M., Koul, R., & Dembowski, J. Computer-Assisted Anomia Treatment for Persons with Chronic Aphasia: Generalization to Untrained Words. Journal of Medical Speech-Language Pathology. 2014; 21(2): 149–163. Full text

Trying it without paying for it

Whether speech recognition suits a particular person is an empirical question, and a week is usually enough to answer it. The first week of Parrot Software is free and does not require a credit card. Use a headset, practise in a quiet room, and see how often it hears correctly. How much practice is enough covers what a realistic schedule looks like afterwards.