By Ian DeNolfo, Executive Director of The Voice Foundation
Executive Overview
We stand at a profound inflection point in the history of voice science—a historic juncture where artificial intelligence (AI) and machine learning (ML) are fundamentally reshaping how researchers conduct investigations, how clinical data is analyzed, and how patients receive specialized care. While global discourse surrounding AI often reads like a chronicle of overnight disruption, the reality within voice science and medicine is far more deeply rooted. Decades before generative pre-trained transformers and conversational agents entered public consciousness, visionary researchers were already harnessing early neural networks to decode the hidden complexities of the human voice.
Today, this digital evolution has transitioned from an experimental sub-discipline into an exponential force. Through its flagship publication, the Journal of Voice, The Voice Foundation has documented this transformation for over thirty years. What began as pioneering queries into spectral pattern recognition in the mid-1990s has erupted into a staggering surge of modern research. In fact, over 63 percent of all AI-related papers published in the Journal of Voice over the last three decades have appeared in just the past three years.
This article explores the historical trajectory of AI in voice science, examines the breakthrough methodologies transforming voice into a powerful digital biomarker for systemic and neurological health, reviews emerging clinical debates regarding conversational AI, and outlines the urgent need for a unified ethical framework as our global community navigates this extraordinary frontier.

A Legacy of Innovation: The Foundations of Interdisciplinary Voice Care
To understand the magnitude of today’s technological leap, one must first appreciate the institutional bedrock upon which it rests. In 1969, Dr. Wilbur James Gould founded The Voice Foundation in New York City during an era when the interdisciplinary care of the human voice was virtually nonexistent. Dr. Gould possessed the groundbreaking foresight to bring together a diverse coalition of physicians, vocal scientists, speech-language pathologists, performing artists, and master teachers. His goal was simple yet revolutionary: to break down academic silos and share specialized knowledge regarding the care and preservation of the professional voice user.
The Foundation established its first Annual Symposium—Care of the Professional Voice—in 1972, followed swiftly by its inaugural Gala (later celebrated as Voices of Summer) in 1973. For more than five decades, the Foundation has served as a vital bridge connecting art and science, the clinical examination room and the performance stage.
Since 1989, The Voice Foundation has operated under the visionary leadership of Dr. Robert Thayer Sataloff, an internationally renowned otolaryngologist, master clinician, professional singer, and conductor. Dr. Sataloff has authored more than 1,200 publications, including 79 authoritative textbooks. Under his stewardship, the Foundation relocated its headquarters to Philadelphia and dramatically expanded its global reach through advanced scientific research and educational initiatives.
Today, the annual Philadelphia Symposium draws hundreds of medical doctors, scientists, academics, speech-language pathologists, and performing artists from every corner of the globe. At the center of this academic ecosystem stands the Journal of Voice, the premier peer-reviewed journal dedicated exclusively to voice science and medicine. It is within the digital and physical pages of this journal that the marriage of artificial intelligence and otolaryngology has been meticulously recorded, debated, and advanced.

Detailed Chronology: Thirty Years of AI Research in Voice Science
Most contemporary observers assume that the integration of artificial intelligence into medicine is a recent phenomenon catalyzed by the arrival of modern cloud computing. Yet, the Journal of Voice has been publishing peer-reviewed AI research for over three decades.
The timeline of machine learning in voice science began in earnest in 1994—thirty-one years ago—when the journal published a landmark paper by Rihkanen and colleagues titled "Spectral Pattern Recognition of Improved Voice Quality." Utilizing early neural networks to analyze voice signals, that paper was published during the exact same calendar year that the World Wide Web was first emerging into public consciousness. Voice scientists were exploring artificial intelligence before the vast majority of professionals had ever sent an electronic mail message.
However, the historical progression of this research reveals a dramatic acceleration curve:
- 1994–2015 (The Incubation Era): 17 papers published over two decades, laying the theoretical and computational groundwork for automated acoustic analysis.
- 2016–2019 (The Methodological Expansion): 27 papers published as deep learning architectures began replacing traditional statistical algorithms.
- 2020–2022 (The Clinical Diversification): 16 papers exploring telemedicine, remote monitoring, and early pathological screenings during global health disruptions.
- 2023–2025 (The Exponential Explosion): 102 papers published in just three years—representing 63 percent of all AI literature in the journal’s history.
To put this trajectory into perspective, the Journal of Voice published 51 AI-related papers in the year 2025 alone. That single year surpassed the total output of the first twenty years of AI research combined. This is no longer gradual academic progress; it is an exponential transformation of an entire scientific discipline.

Supporting Context and Metrics: Real-World Impact and Global Reach
Academic publishing is often evaluated through the cold lens of citation metrics, but the true measure of scientific literature lies in its practical application by clinicians, therapists, and researchers on the front lines of healthcare.
Usage statistics for AI-related articles within the Journal of Voice demonstrate extraordinary global engagement. The average AI paper published in the journal has been downloaded nearly 1,000 times—a remarkable figure reflecting intense professional interest. Furthermore, specific breakthrough papers have shattered standard academic readership ceilings:
- COVID-19 Detection Study: A machine learning study examining acoustic indicators for respiratory illness has been downloaded over 7,200 times, reaching eleven times the median readership for articles in its issue.
- Deep Learning for Pathology (Fang et al., 2019): This foundational study on cepstrum vectors amassed 193 citations and nearly 5,000 downloads, cementing itself as essential reading for computational otolaryngology.
- Machine Learning Survey (Hegde et al.): Providing an exhaustive landscape map of voice disorder detection algorithms, this survey accumulated 147 citations and close to 5,000 downloads.
These numbers do not merely represent passive citation harvesting. They represent clinicians, speech-language pathologists, and biomedical engineers actively seeking out actionable methodologies to improve patient diagnostics and treatment outcomes.
A Global Research Community
This scientific revolution is powered by a truly international network of scholars. Leading contributors to the journal’s AI corpus include Jérôme René Lechien (6 papers), Dimitar Deliyski (5 papers), Stephanie Zacharias (5 papers), and Ahmed Yousef (4 papers as first author), alongside prominent international investigators such as Paavo Alku, Leonardo Wanderley Lopes, and Maryam Naghibolhosseini. Their collaborative efforts span institutions across every inhabited continent, unified by a shared commitment to decoding the human voice through computation.

Landmark Research: Reshaping Clinical Diagnostics
Deep Learning and Voice Pathology Detection
The journey from simple spectral pattern recognition to sophisticated deep neural networks has unlocked diagnostic capabilities previously thought impossible. In their 2019 study, "Detection of Pathological Voice Using Cepstrum Vectors: A Deep Learning Approach," Fang and colleagues proved that deep neural networks could identify vocal fold pathologies with staggering accuracy using subtle acoustic features imperceptible to the human ear.
This paved the way for subsequent innovations, such as Chen and Chen’s work on deep neural networks for voice classification, Fujimura’s development of one-dimensional convolutional neural networks, and Cho and Choi’s comparative analysis of CNN models applied to laryngoscopic imaging. Each successive study built upon its predecessor, pushing the boundaries of automated screening.
The Voice as a Digital Biomarker
Perhaps the most paradigm-shifting concept to emerge from recent literature is the characterization of the human voice as a digital biomarker—a non-invasive, continuous window into systemic and neurological health.
The Journal of Voice has published groundbreaking investigations into using automated voice analysis to detect and monitor neurodegenerative conditions like Parkinson’s disease. Notably, Hemmerling and Wójcik-Pędziwiatr’s 2022 paper on predicting Parkinson’s severity from voice signals quickly garnered dozens of citations, pushing the field closer to routine clinical integration.

Beyond neurology, systematic reviews published in the journal have begun exploring voice quality as a digital biomarker for affective disorders, including clinical depression and bipolar disorder. The societal implications are profound: routine, automated voice screenings during standard telehealth check-ups could enable earlier mental health interventions, substantially reducing the crushing burden on overstretched healthcare infrastructure.
AI Chatbots in Clinical Decision-Making
As conversational artificial intelligence matured, the journal expanded its scope to examine the role of large language models (LLMs) in direct clinical workflows. A prime example is the 2025 paper by Dronkers and colleagues, "Evaluating the Potential of AI Chatbots in Treatment Decision-making for Acquired Bilateral Vocal Fold Paralysis," which immediately sparked intense scholarly debate, multiple letters to the editor, and subsequent investigations into tools like ChatGPT-4o’s accuracy in evaluating laryngeal images. The medical community is actively and transparently grappling with the capabilities and limitations of conversational AI in real-time clinical environments.
Expert Perspectives on AI in Voice Science
To contextualize these technological shifts, we turn to two leading voices within our community who offer critical insights into the computational and institutional futures of voice science.
Dr. Mark Berardi: AI as a Tool for Complexity
Dr. Mark Berardi bridges the worlds of physics, computation, and voice science. He views artificial intelligence not as an autonomous oracle, but as an essential instrument for managing biological complexity.

"Machine learning and AI are tools to help with complexity, and communication is incredibly complex in its neurobiological and physiological processes. So I think the application is warranted," Dr. Berardi explains.
His current research focuses on extracting health insights—such as markers for aging and depression—from easily acquired speech signals. While he acknowledges that generative AI has not yet completely revolutionized raw discovery, he highlights its immediate utility in eliminating research bottlenecks: "I can now quickly create bespoke code and edit it with natural language prompts."
Intriguingly, Dr. Berardi is also pioneering research into human-AI communication itself. By studying how individuals alter their linguistic patterns when speaking to chatbots versus humans—or how "synthetic" communication formats like Zoom alter vocal delivery—he suggests that voice science must soon expand its purview. Researchers must study not only the human voice in isolation, but the human voice in active conversation with machines.
Dr. Eric Hunter: Embracing Change with Ethical Clarity
Dr. Eric Hunter focuses heavily on the institutional and academic implications of widespread AI adoption. He issues a clear and pragmatic warning to scholars and educators:

"These tools aren’t just novelties. They are rapidly becoming integrated into how we conduct literature reviews, organize data, extract themes from large bodies of text, and even summarize findings."
Dr. Hunter stresses that large language models and automated platforms are permanently embedded within academic workflows, from manuscript preparation to peer review. Rather than resisting this inevitable shift, he advocates for proactive adaptation centered around four core ethical principles:
- Absolute Transparency: Clear disclosure of where and how AI tools were utilized in the research and writing process.
- Human Accountability: Uncompromising insistence that human authors remain entirely responsible for the factual accuracy and integrity of their work.
- Data Privacy and Security: Rigorous protection of sensitive patient voice data when utilizing third-party computational engines.
- Editorial Integrity: Establishing robust guidelines for peer reviewers to prevent the unverified outsourcing of manuscript evaluation to automated systems.
"Our field will benefit most," Dr. Hunter concludes, "if we embrace the productivity these tools offer while also building a shared ethical framework for their responsible use."
Future Outlook: The Path Forward
As the Journal of Voice enters a new decade of digital transformation, it remains resolutely committed to publishing rigorous, peer-reviewed research that expands scientific horizons while upholding the highest standards of academic integrity. However, the road ahead demands concerted, collective action across three distinct pillars:

- Thoughtful Engagement with Complexity: Researchers and clinicians must master AI tools to amplify their own expertise rather than abdicate clinical judgment to algorithms. Understanding tool limitations is just as vital as celebrating computational power.
- Collaborative Ethical Governance: Dr. Hunter’s call for a shared ethical framework cannot be answered by a single institution or journal. It requires a united, international front comprising otolaryngologists, speech-language pathologists, ethicists, and engineers.
- Continued Open Science: The 161 AI papers published in our journal thus far represent merely the foundation. Every investigator, clinician, and educator harbors valuable insights that can propel our collective understanding forward.
An Extraordinary Moment in Human History
For hundreds of thousands of years, the human voice has served as our primary instrument of connection, emotional expression, and personal identity. It carries our health, our history, and our very essence in ways that no other biological signal can match.
Now, for the first time in human history, we possess computational tools sophisticated enough to truly decode that complexity—to uncover what the voice reveals about our brains, our bodies, and our overall well-being. We have at our disposal technologies that can extend the reach of expert clinicians, detect subtle pathology long before symptoms manifest, and democratize access to high-quality voice care for vulnerable populations across the globe.
We are living through an extraordinary moment. The Journal of Voice and The Foundation established by Dr. Gould stand proudly at the center of this revolution. The next decade promises discoveries that will dwarf the achievements of the past thirty years, and we look forward to sharing every breakthrough with our global community.
About the Author
Ian DeNolfo is Executive Director of The Voice Foundation, which publishes the peer-reviewed Journal of Voice. A graduate of The Juilliard School and The Curtis Institute of Music, he formerly performed as a leading tenor at major opera houses worldwide before transitioning his professional expertise to health science administration and leadership at The Voice Foundation.
