By Ian DeNolfo
Executive Director, The Voice Foundation
Executive Overview
We stand at a profound inflection point in the scientific study of the human voice. Across research laboratories, academic medical centers, and clinical practices worldwide, artificial intelligence (AI) is fundamentally reshaping how we conduct research, analyze acoustic datasets, and deliver patient care.
This digital renaissance does not represent an abrupt departure from our past; rather, it is the acceleration of a trajectory decades in the making. The Journal of Voice—the premier peer-reviewed publication dedicated to voice science and medicine—has been quietly pioneering the integration of machine learning and artificial intelligence for over thirty years. Long before mainstream awareness of neural networks or automated data processing, our field recognized that computational tools would be essential to unlocking the vast, intricate complexities of human vocal physiology and pathology.
Today, this quiet integration has exploded into exponential growth. Out of 161 AI-focused papers published in the Journal of Voice over its history, an astonishing 63 percent (102 papers) have been published in just the last three years. In 2025 alone, the journal has released 51 AI-related studies—surpassing the total output of the publication’s first twenty years combined.
This is not merely incremental academic progress. It represents a paradigm shift. By treating the human voice as a rich, non-invasive digital biomarker, researchers and clinicians are now detecting neurological disorders, mental health fluctuations, and respiratory pathologies well before clinical symptoms manifest on the surface. As we navigate this new frontier, balancing technological innovation with rigorous ethical frameworks will determine how successfully we translate these computational breakthroughs into better patient outcomes.

A Legacy of Interdisciplinary Innovation
To understand the magnitude of today’s technological transformation, one must look to the foundational history of The Voice Foundation. In 1969, Dr. Wilbur James Gould founded the organization in New York City at a time when comprehensive, interdisciplinary care for the human voice was virtually nonexistent.
Dr. Gould possessed the visionary foresight to break down traditional academic and clinical silos. He brought together otolaryngologists, vocal scientists, speech-language pathologists, performing artists, and master voice teachers to share expertise in the preservation and treatment of the professional voice. In 1972, the Foundation hosted its inaugural Annual Symposium—Care of the Professional Voice—establishing a global academic anchor that continues to draw hundreds of international medical and artistic professionals to Philadelphia every year.
Since 1989, The Voice Foundation has been guided by Dr. Robert Thayer Sataloff, an internationally renowned otolaryngologist, professional singer, and conductor. Dr. Sataloff has authored more than 1,200 publications, including 79 textbooks. Under his stewardship, the Foundation relocated to Philadelphia and expanded its global reach, cementing the Journal of Voice as the definitive authority on voice science and medicine.
For over five decades, the Foundation has built bridges between art and science, and between the clinic and the stage. Today, those bridges are reinforced with algorithms, deep learning models, and neural architectures.
Detailed Chronology: 30 Years of AI Research in Voice Science
A common misconception in modern academic discourse is that artificial intelligence in medicine is an overnight phenomenon born from the recent explosion of large language models. In voice science, however, the digital foundation was laid more than three decades ago.

The Journal of Voice published its first AI-centric paper in 1994—thirty-one years ago. Titled "Spectral Pattern Recognition of Improved Voice Quality" by Rihkanen and colleagues, the study deployed foundational neural networks to analyze voice metrics. To put this milestone into historical perspective, 1994 was the exact year the World Wide Web was beginning to emerge into public consciousness. Voice scientists were utilizing neural networks to evaluate acoustic data before the vast majority of professionals had ever sent an electronic mail message.
The Publication Timeline: An Exponential Curve
The historical evolution of AI literature within the Journal of Voice illustrates a dramatic shift from exploratory curiosity to widespread clinical application:
- 1994 – 2015: 17 papers. During this foundational era, computational models were largely experimental, serving as proof-of-concept tests for automated acoustic classification.
- 2016 – 2019: 27 papers. As machine learning matured and computational power expanded, researchers began exploring advanced algorithms for automated pathology detection.
- 2020 – 2022: 16 papers. Spanning the height of the COVID-19 pandemic, this period saw accelerated interest in remote, non-contact biometric screening tools.
- 2023 – 2025: 102 papers. An explosive surge driven by deep learning architectures, cloud computing accessibility, and the maturation of generative AI tools.
In the year 2025 alone, the journal published 51 AI-related papers. This staggering vertical trajectory signals that artificial intelligence is no longer a peripheral sub-specialty; it has become the central engine driving modern voice research.
Supporting Context & Metrics: Real-World Impact and Global Reach
Academic research is frequently evaluated through citation metrics, but the true measure of scientific literature lies in its practical application by frontline practitioners. The AI research published in the Journal of Voice is actively consumed, downloaded, and implemented by clinicians across the globe.
Access and Engagement Metrics
- High Readership: The average AI-focused paper published in the journal has been downloaded nearly 1,000 times.
- The COVID-19 Benchmark: The most-accessed AI paper in the journal’s history—a pivotal study exploring machine learning applications for COVID-19 detection—has amassed over 7,200 downloads, eclipsing the median download rate for its issue by a factor of eleven.
- Citation Leaders: The two most-cited papers—Fang and colleagues’ deep learning study (193 citations) and Hegde’s machine learning survey (147 citations)—have each been downloaded nearly 5,000 times, ranking consistently among the top-accessed articles in their respective volumes.
A Global Community of Scholars
This digital transformation is driven by a vast, interconnected international research community. Key contributors shaping the pages of the Journal of Voice include:

- Dr. Jérôme René Lechien (6 papers)
- Dr. Dimitar Deliyski (5 papers)
- Dr. Stephanie Zacharias (5 papers)
- Dr. Ahmed Yousef (4 papers as first author)
- Dr. Paavo Alku, Dr. Leonardo Wanderley Lopes, and Dr. Maryam Naghibolhosseini (4 papers each)
Supported by dozens of interdisciplinary research teams spanning every continent, these scholars are proving that the computational revolution in voice science is a truly global endeavor.
Landmark Research Reshaping the Field
Deep Learning for Voice Pathology Detection
In 2019, Fang and colleagues published "Detection of Pathological Voice Using Cepstrum Vectors: A Deep Learning Approach," which has since become the most-cited AI paper in the journal’s history (193 citations). This landmark study demonstrated that deep neural networks could identify pathological conditions in the human voice with extraordinary accuracy—leveraging acoustic features far too subtle for the unassisted human ear to perceive.
This foundational work built upon earlier mapping efforts, such as Hegde and colleagues’ comprehensive "Survey on Machine Learning Approaches for Automatic Detection of Voice Disorders" (147 citations), alongside pivotal 2017 studies by Al-Nasheri et al. examining Multidimensional Voice Program parameters and correlation functions. More recently, researchers like Chen and Chen (2022) with deep neural networks for voice classification, Fujimura with one-dimensional convolutional neural networks, and Cho and Choi’s comparative CNN analysis of laryngoscopic images have continuously pushed the boundaries of diagnostic precision.
Voice as a Digital Biomarker
Perhaps the most medically transformative concept to emerge from recent literature is the characterization of the voice as a digital biomarker—a non-invasive window into systemic, whole-body health.
Significant breakthroughs have been made in utilizing automated acoustic analysis to detect and monitor neurodegenerative conditions. For example, Hemmerling and Wójcik-Pędziwiatr’s 2022 paper on predicting the severity of Parkinson’s disease from voice signals has accrued 29 citations in just three years, paving the way for continuous, remote patient monitoring.

Beyond neurology, systematic reviews now highlight voice quality as an emerging digital biomarker for psychiatric conditions such as clinical depression and bipolar disorder. The clinical implications are immense: routine, automated voice screening could enable early psychiatric intervention, alleviate the immense burden on overstretched mental health systems, and provide objective metrics for therapeutic efficacy.
AI Chatbots in Clinical Practice
The year 2025 has introduced a new wave of scholarship examining the integration of conversational AI and large language models into clinical decision-making. Dronkers and colleagues’ paper, "Evaluating the Potential of AI Chatbots in Treatment Decision-making for Acquired Bilateral Vocal Fold Paralysis," captured immediate scholarly attention, accumulating 18 citations within months of publication.
Spurring energetic letters to the editor and follow-up studies evaluating models like ChatGPT-4o in analyzing complex laryngeal images, the scientific community is actively grappling with the realities of conversational AI in medical settings in real time.
Official Statements and Expert Perspectives
To capture the strategic and philosophical dimensions of this shift, we turn to two leading voices in our academic community: Dr. Mark Berardi and Dr. Hunter.
Dr. Mark Berardi: AI as a Tool for Complexity
Drawing from a rich background in physics and computation, Dr. Mark Berardi approaches voice science through the lens of complex systems.

"Machine learning and AI are tools to help with complexity, and communication is incredibly complex in its neurobiological and physiological processes. So I think the application is warranted," Dr. Berardi notes.
His current research targets voice-based digital biomarkers for aging and depression. While noting that acquiring speech and voice signals is technically straightforward today, the true hurdle remains the inherent complexity of human communication itself. Furthermore, Dr. Berardi offers a pragmatic perspective on generative AI:
"While it hasn’t yet created a major change to research methodologies on its own, it excels at addressing computational bottlenecks—particularly in custom coding and data processing. I can now quickly create bespoke code and edit it using natural language prompts."
Intriguingly, Dr. Berardi is also pioneering research into human-AI communication: how humans unconsciously adapt their speech patterns when interacting with chatbots versus living interlocutors, and how "synthetic" communication environments (such as Zoom calls) alter interpersonal linguistics. Understanding the human voice in conversation with machines may soon become a vital sub-discipline of voice science.
Dr. Eric Hunter: Embracing Change with Ethical Clarity
Dr. Eric Hunter emphasizes the institutional and administrative imperatives of adopting artificial intelligence within academic and clinical workflows.

"These tools aren’t just novelties," Dr. Hunter observes. "They are rapidly becoming integrated into how we conduct literature reviews, organize data, extract themes from large bodies of text, and even summarize findings."
Dr. Hunter issues an inescapable conclusion: large language models and AI-driven platforms are permanent fixtures of modern scholarship. Rather than resisting this technological wave, the academic community must establish proactive guidelines. He outlines four essential pillars for responsible AI integration in research and publishing:
- Transparency: Authors must explicitly disclose if and how generative tools were utilized during manuscript preparation, data analysis, or literature synthesis.
- Accountability: Human researchers remain entirely responsible for the factual accuracy, integrity, and ethical soundness of published work.
- Editorial Oversight: Reviewers and journal editors must be equipped with clear policies to evaluate AI-assisted submissions fairly and rigorously.
- Data Privacy: Strict protocols must govern the handling of patient voice data to protect personal health information in machine learning pipelines.
"Our field will benefit most," Dr. Hunter concludes, "if we embrace the productivity these tools offer while also building a shared ethical framework for their responsible use."
Future Outlook and the Path Forward
As the Journal of Voice charts its course into the coming decade, it will remain the premier global platform for this rapidly evolving discipline. However, realizing the full potential of AI-driven voice science requires deliberate, collective action across three key fronts:
- First, thoughtful engagement: We must view AI not as a replacement for human clinical judgment, but as an amplifier of human expertise. Clinicians and researchers must understand both the immense capabilities and the inherent limitations of algorithmic tools.
- Second, collaborative ethics: Dr. Hunter’s call for a shared ethical framework cannot be answered by isolated institutions. Journal editors, academic societies, clinicians, and tech developers must forge universal standards together.
- Third, continued contribution: The 161 AI papers published in our journal to date represent merely the dawn of this movement. Every voice professional holds unique clinical insights that can enrich our collective understanding.
An Extraordinary Moment
For hundreds of thousands of years, the human voice has served as our ultimate instrument of connection, emotional expression, and personal identity. It transmits our inner states, our psychological wellbeing, and our physical health in ways that no other biological signal can match.

Now, for the first time in human history, we possess technological tools sophisticated enough to truly decode that complexity—tools capable of extracting health insights hidden deep within acoustic waves, extending the reach of expert clinicians, and democratizing access to specialized voice care across the globe.
We are living through an extraordinary moment in medical and scientific history, and The Voice Foundation and the Journal of Voice stand proudly at its epicenter. The next decade promises discoveries we are only beginning to imagine. We look forward to sharing them with you.
About the Author
Ian DeNolfo is Executive Director of The Voice Foundation, which publishes the Journal of Voice. A graduate of The Juilliard School and The Curtis Institute of Music, he formerly performed as a leading tenor at opera houses worldwide before transitioning to arts administration and scientific leadership.
