In This Section
91视频 Cognitive Psychologist鈥檚 Work Has Generated Over 14,000 Language Studies
TalkBank began as a project of convenience in the 1980s, but has grown into a family of databases that has transformed the field of language science.
By Jason Bittel Email Jason Bittel
Most people have never heard of . But for those working in the language sciences, the family of 14 databases created by 91视频鈥檚 Brian MacWhinney has become the foundation for research into everything from how children learn to speak and how people pick up a second language to patterns that can reveal the onset of dementia.
At their heart, TalkBank and its spinoffs 鈥 including AphasiaBank and DementiaBank 鈥 are enormous, standardized databases of audio and video recordings of people talking that allow scientists to study many different components of language.
Before these databases, scientists relied on their own, often highly personalized transcription formats, codes and language analysis programs. And like the aftermath of the Tower of Babel, this severely limited how much cooperation researchers working in any given lab could offer one another.
But now, after nearly five decades, the TalkBank Project is the world鈥檚 largest open-access integrated repository for spoken-language data. It has ushered in a revolution to the language science field and yielded more than 14,000 scientific publications.
鈥淲e have people all over the world, with enormous amounts of data, and everybody鈥檚 using it,鈥 said Brian MacWhinney, Teresa Heinz Professor of Cognitive Psychology in 91视频鈥檚 Dietrich College of Humanities and Social Sciences and creator of TalkBank.
In fact, this work has been so transformative, MacWhinney was recently awarded the , a lifetime achievement award presented by the Faculty of Humanities of The Hong Kong Polytechnic University (PolyU).
The prize recognizes MacWhinney鈥檚 鈥渓ifetime of distinguished contributions to language science, encompassing integrative theoretical innovation, research infrastructure development and lasting international impact on the study of human language,鈥 according to PolyU.
鈥淏rian鈥檚 work is the gift that keeps on giving,鈥 said Susanne Ferber, professor and head of the Department of Psychology. 鈥淗is foundational contributions to science go well beyond understanding the complexities of first or second language acquisition to advance our insights into how spoken language offers a window into brain health.鈥
The Power of Standardization
MacWhinney still recalls the moment in 1981 when he and his early colleagues started down the path that he鈥檇 blaze for the next half-century.
鈥淚 can remember very clearly that we were at a meeting in Nijmegen, and we had some child language transcripts that had been mimeographed,鈥 said MacWhinney, referencing the predecessor technology to the copier. 鈥淎nd we were marking them up with comments on the side, in pencil, and it occurred to me鈥 they had just come out with the IBM PC. Why not make files that could be more quickly and easily circulated?鈥
By 1984, MacWhinney and his team had won the support of the MacArthur Foundation, which enabled them to bring 20 of the world鈥檚 top language researchers to Concord, Massachusetts, to hammer out the details of what would soon be known as the Child Language Data Exchange System, or CHILDES.
For the first time, CHILDES provided language researchers across the world with standardized data. But it wasn鈥檛 long before MacWhinney saw the need to build off of CHILDES鈥 success to develop still more databases that could be used for adults and other populations.
It was in 2001, with the support of a National Science Foundation Infrastructure Grant, that MacWhinney and his team created a standardized transcription and coding system, called CHAT, as well as a standardized analysis program, CLAN. Combined with the database, CHAT and CLAN made up the skeleton of TalkBank 颅颅鈥 a formula the team would repeat to create 14 more standardized databases that have become invaluable in the language sciences and which can be used across 18 languages.
鈥淭here鈥檚 a database for dementia, called DementiaBank. There鈥檚 a database for autism, called ASD Bank. There鈥檚 traumatic brain injury, there鈥檚 stuttering, there鈥檚 second language learning,鈥 said MacWhinney.
Name a topic of interest in language science, and there鈥檚 now a standardized database for it.
What鈥檚 Next for Language Science?
As just one recent example of the advances being made in language science as a result of MacWhinney鈥檚 work, he pointed to a recent effort to develop better automatic speech recognition for children, who have distinct vocal characteristics, inconsistent pronunciation and who are still learning how to control their mouths when they speak.
Known as the , over 800 language scientists submitted some 2,100 solutions to the problem prompts, which made use of data from the CHILDES and PhonBank databases in TalkBank. In the end, the winners blew MacWhinney away.
鈥淭he top solvers cut the error rate of the best existing children鈥檚 speech model by more than 40%,鈥 he said. 鈥淭hat kind of increase in accuracy is incredible. In computer science, you might shoot for 1% or 2%.鈥
Similarly, MacWhinney believes there is enormous promise in the work related to speech patterns and dementia.
鈥淚t鈥檚 one area where we鈥檝e only just started about five years ago, and it鈥檚 really the only publicly available set of data on language and dementia,鈥 he said. 鈥淪everal hundred computer science labs, as well as private companies, are now using it to develop methods for early detection.鈥
鈥淭he best part about it is you can do it without putting somebody in an MRI machine or extracting fluids from their spinal cord,鈥 said MacWhinney. 鈥淵ou just take a measure of their speech, and you can really detect onset of dementia with high accuracy. It鈥檚 not perfect, but we鈥檙e talking up to 95% accuracy.鈥
Of course, TalkBank is also a natural partner for the fields of machine learning and artificial intelligence, including efforts to get AI models to learn language like children called BabyLLM.
鈥淎I and machine learning need lots of data, and TalkBank provides it,鈥 said MacWhinney. 鈥淎t the same time, AI is reshaping TalkBank data, analysis and infrastructure.鈥
Even though updating and managing the databases now takes up so much of MacWhinney鈥檚 time that he can no longer do the experimental or theoretical work that got him into the field, he loves what he does.
鈥淚鈥檓 not retiring, because it鈥檚 fun,鈥 said MacWhinney, who is fluent in Spanish, Hungarian, German and French. 鈥淚t鈥檚 fun to look at language.鈥