By Ian DeNolfo
Executive Director, The Voice Foundation
Executive Overview
We stand at a definitive inflection point in the scientific study of human vocalization. For over half a century, the intersection of voice care, clinical otolaryngology, and performing arts has relied on human intuition, physiological measurements, and acoustic observation. Today, that paradigm is undergoing a fundamental restructuring. Artificial intelligence (AI) and machine learning (ML) are no longer futuristic concepts hovering on the horizon of biomedical engineering; they are active, indispensable forces reshaping how we conduct research, decode complex acoustic datasets, and deliver patient care.
This technological renaissance does not mark a departure from the historical foundations of voice science. Rather, it represents the amplification of a legacy built over five decades. Through the pages of the Journal of Voice—the premier peer-reviewed publication dedicated to voice science and medicine—researchers have quietly tracked, shaped, and accelerated this computational revolution.
As we analyze the publication metrics and clinical breakthroughs of recent years, a striking narrative emerges: over 60 percent of all AI-related papers published in the journal’s history have appeared in just the last three years. This is not gradual, linear growth. It is an exponential transformation, cementing voice as a premier digital biomarker for systemic human health, and redefining the boundaries of what is medically possible.

A Legacy of Interdisciplinary Innovation
To understand the magnitude of today’s technological integration, one must examine the institutional bedrock upon which modern voice science was built. In 1969, Dr. Wilbur James Gould founded The Voice Foundation in New York City. At a time when interdisciplinary care for the human voice was virtually non-existent, Dr. Gould possessed the groundbreaking foresight to bring together a disparate group of professionals: physicians, scientists, speech-language pathologists, performing artists, and vocal pedagogues. His mission was simple yet radical—to foster a collaborative environment where cross-disciplinary expertise could be pooled for the betterment of the professional voice user.
The momentum generated by Dr. Gould’s vision quickly materialized into institutional milestones. In 1972, the Foundation hosted its first Annual Symposium: Care of the Professional Voice, followed a year later by its inaugural Gala (later renamed Voices of Summer). For more than fifty years, these gatherings have served as the ultimate bridge between art and science, connecting the operating room, the scientific laboratory, and the theatrical stage.
Since 1989, The Voice Foundation has been guided by Dr. Robert Thayer Sataloff, an internationally renowned otolaryngologist, professional singer, and conductor. Dr. Sataloff’s staggering academic output includes more than 1,200 publications and 79 textbooks. Under his stewardship, the Foundation relocated its headquarters to Philadelphia and expanded its global footprint, cementing its status as the world’s leading authority on voice education and interdisciplinary research. Today, the annual Philadelphia Symposium draws hundreds of medical, scientific, academic, and artistic minds from across the globe, while the Journal of Voice continues to set the gold standard for peer-reviewed literature in the discipline.
Detailed Chronology: Thirty Years of AI Research in Voice Science
While mainstream society has only recently grown consumed by the rapid ascent of generative AI, the academic community surrounding the Journal of Voice was engaging with computational neural networks decades earlier.

The historical timeline of machine learning in voice science is longer and deeper than most realize. In 1994—thirty-one years ago, during the absolute infancy of the World Wide Web, when sending an email was still a novelty for the general public—the Journal of Voice published a pioneering paper by Rihkanen and colleagues titled "Spectral Pattern Recognition of Improved Voice Quality." This foundational study utilized early neural networks to analyze voice characteristics, proving that machine-based pattern recognition could successfully quantify alterations in vocal quality.
For many years, computational applications in voice science progressed incrementally. Between 1994 and 2015, the journal published 17 AI-related papers. The pace quickened slightly between 2016 and 2019, yielding 27 papers as machine learning models matured. Between 2020 and 2022, despite global disruptions, 16 papers were added to the canon.
Then came the inflection point. Between 2023 and 2025, a staggering 102 AI-related papers were published in the Journal of Voice, accounting for 63 percent of all computational studies in the journal’s history. In the single year of 2025, the journal published 51 AI-related articles—surpassing the total output of the first twenty years of AI publishing combined. This explosive trajectory signals that artificial intelligence has transitioned from a peripheral computational tool to the core engine driving modern voice research.
Supporting Context and Metrics: Real-World Impact and Landmark Studies
The value of academic research is measured not merely by its volume, but by its reach, replication, and practical application. The AI-related literature published in the Journal of Voice has achieved extraordinary penetration across clinical and computational communities.

On average, an AI-focused paper published in the journal is downloaded nearly 1,000 times—a testament to the high demand for actionable computational methodologies among clinicians. The real-world impact is further illustrated by landmark studies that have captured the global academic imagination:
- Machine Learning for COVID-19 Detection: The single most-accessed AI paper in the journal’s history—a machine learning study focused on utilizing voice characteristics to detect COVID-19—has amassed over 7,200 downloads, representing eleven times the median readership for articles in its respective issue.
- Deep Learning for Pathology Detection: Fang and colleagues’ 2019 study, "Detection of Pathological Voice Using Cepstrum Vectors: A Deep Learning Approach," stands as the most-cited AI paper in the journal’s history with 193 citations. This research proved that deep neural networks could detect voice pathologies with astonishing precision using acoustic features entirely imperceptible to the human ear.
- Comprehensive Methodological Frameworks: Hegde and colleagues’ comprehensive "Survey on Machine Learning Approaches for Automatic Detection of Voice Disorders" (147 citations) has become mandatory reading for engineers and clinicians entering the field, mapping out the architecture of machine learning algorithms tailored for vocal diagnostics. Additional high-impact contributions from Al-Nasheri et al. (2017) explored Multidimensional Voice Program parameters and correlation functions, securing nearly 100 citations each.
- Conversational AI in Clinical Settings: Pushing into the present day, 2025 has seen an influx of papers examining conversational AI in clinical decision-making. Dronkers and colleagues’ paper, "Evaluating the Potential of AI Chatbots in Treatment Decision-making for Acquired Bilateral Vocal Fold Paralysis," rapidly accumulated 18 citations within its first few months of publication, sparking rigorous debate and subsequent letters to the editor regarding tools like ChatGPT-4o’s efficacy in analyzing laryngeal imagery.
Voice as a Digital Biomarker
Perhaps the most philosophically profound shift in modern voice science is the recognition of the human voice as a digital biomarker—a non-invasive, accessible window into systemic neurological, psychological, and respiratory health.
Recent literature highlights the utilization of voice analysis to detect and monitor neurodegenerative conditions such as Parkinson’s disease. For instance, Hemmerling and Wójcik-Pędziwiatr’s 2022 paper predicting Parkinson’s severity from voice signals quickly gathered significant citations, steering the field closer to routine clinical implementation. Furthermore, systematic reviews exploring vocal quality as a digital biomarker for clinical depression and bipolar disorder open up radical possibilities: the continuous, passive monitoring of mental health states through routine vocal interactions, facilitating early therapeutic intervention and alleviating the strain on overburdened healthcare systems.
Official Statements and Expert Perspectives
To contextualize this computational revolution, prominent thought leaders within the Voice Foundation ecosystem offer vital insights regarding the opportunities and ethical obligations of integrating AI into scientific workflows.

Dr. Mark Berardi: Managing Complexity and Synthetic Communication
Dr. Mark Berardi, whose academic background spans physics and computation, approaches voice science through the lens of complex systems.
"Machine learning and AI are tools to help with complexity, and communication is incredibly complex in its neurobiological and physiological processes," Dr. Berardi notes. "So I think the application is warranted."
Focusing his research on voice-based digital biomarkers for aging and depression, Dr. Berardi emphasizes that while acquiring vocal signals has become frictionless, interpreting the underlying biological mechanics remains challenging—an obstacle tailor-made for AI’s pattern-recognition capabilities.
On generative AI, Dr. Berardi maintains a pragmatic view. While it may not have entirely rewritten the foundational theories of biology overnight, it has shattered traditional research bottlenecks. "I can now quickly create bespoke code and edit it with natural language prompts," he explains.

Crucially, Dr. Berardi’s work also examines the emerging phenomenon of human-AI communication. As society shifts toward virtual interactions—ranging from Zoom conferences to conversational chatbots—linguistic and vocal adaptations are occurring in real time. "We are already seeing linguistic differences in chatbot interactions," he observes. The scientific community must now expand its scope to analyze not just the human voice in isolation, but the human voice in dialogue with synthetic intelligence.
Dr. Eric Hunter: Embracing Change with Ethical Clarity
Dr. Eric Hunter focuses heavily on the structural and institutional implications of AI adoption within academia and clinical practice.
"These tools aren’t just novelties," Dr. Hunter warns. "They are rapidly becoming integrated into how we conduct literature reviews, organize data, extract themes from large bodies of text, and even summarize findings."
Acknowledging that large language models (LLMs) and AI-driven platforms are permanently embedded within scholarly infrastructure—from Microsoft Office integrations to automated peer-review assistants—Dr. Hunter issues a definitive call to action. Rather than resisting technological progression, the academic community must establish transparent ethical parameters.

Dr. Hunter outlines four essential pillars for responsible integration:
- Absolute Transparency: Full disclosure of any AI tools utilized in drafting, coding, or data synthesis.
- Human Accountability: Maintaining strict human oversight; algorithms can assist in processing complexity, but human judgment remains the ultimate arbiter of scientific truth.
- Data Integrity & Privacy: Safeguarding patient data and vocal biometric datasets against unauthorized exploitation.
- Equitable Access: Ensuring that AI-driven diagnostic tools are developed and distributed inclusively across diverse global populations.
"Our field will benefit most," Dr. Hunter concludes, "if we embrace the productivity these tools offer while also building a shared ethical framework for their responsible use."
A Global Research Community
The rapid acceleration of AI in voice science is not the isolated achievement of a single laboratory or institution. It is powered by a vast, interconnected global network of researchers spanning every continent.
Key contributors driving the publication metrics of the Journal of Voice include scholars such as Jérôme René Lechien (with 6 papers), Dimitar Deliyski (5 papers), Stephanie Zacharias (5 papers), and Ahmed Yousef (4 papers as first author), alongside international luminaries like Paavo Alku, Leonardo Wanderley Lopes, and Maryam Naghibolhosseini. Supported by dozens of research groups worldwide, this community is actively bridging the gap between computational science and clinical otolaryngology.

Future Outlook and the Path Forward
As the Journal of Voice charts its course into the coming decade, its commitment to rigorous, high-integrity scientific publishing remains unwavering. However, navigating this new frontier requires deliberate, collective action from the global voice care community:
- Mindful Engagement: Researchers and clinicians must approach AI as an instrument for managing complexity—amplifying human expertise rather than abdicating clinical judgment.
- Collaborative Ethics: The establishment of universal ethical guidelines for AI in medical literature and diagnostics requires multidisciplinary cooperation across institutions, journals, and professional societies.
- Continuous Discovery: With 161 AI papers representing merely the foundational chapter of this movement, academics and clinicians are encouraged to share their insights, pushing the boundaries of what vocal biometrics can achieve.
An Extraordinary Moment
For hundreds of thousands of years, the human voice has served as our primary instrument of connection, emotional expression, and identity. It encodes our health, our vitality, and our deepest psychological states in a manner matched by no other biological signal.
Today, for the first time in human history, we possess technological tools sophisticated enough to genuinely decode that complexity. We stand on the verge of empowering clinicians with predictive diagnostics, detecting life-altering pathologies before symptoms manifest, and democratizing access to premier voice care on a global scale.
This is an extraordinary moment in the history of medicine and art. As The Voice Foundation and the Journal of Voice stand firmly at the center of this revolution, the next decade promises discoveries that will redefine not only how we care for the voice, but how we understand the human condition itself.
