Executive Overview
We stand at a profound inflection point in the trajectory of voice science and otolaryngology. For hundreds of thousands of years, the human voice has served as our most fundamental medium of emotional expression, social connection, and personal identity. It carries the distinct cadence of our humanity, reflecting our deepest thoughts, our shifting moods, and the silent, ongoing physiological states of our physical bodies. Yet, for all its evolutionary and communicative importance, the intricate neurobiological and acoustic mechanics of the voice have long eluded complete human comprehension due to their staggering multidimensional complexity.
Today, that paradigm is shifting entirely. The intersection of artificial intelligence (AI), machine learning (ML), and voice science is no longer a speculative frontier of science fiction; it is an active, rapidly accelerating reality. Artificial intelligence is fundamentally reshaping how researchers conduct investigations, how clinical data is analyzed, and how patients receive interdisciplinary care. At the epicenter of this revolution is The Voice Foundation and its premier peer-reviewed publication, the Journal of Voice.
A historical review of the field reveals a surprising truth: the integration of artificial intelligence into voice science is not a sudden, knee-jerk reaction to the recent generative AI boom. Rather, it represents the maturation of over three decades of quiet, pioneering computational research. With more than 63 percent of all AI-related papers in the Journal of Voice published within the last three years alone, the field is undergoing a massive, exponential transformation. As voice is increasingly recognized as a non-invasive digital biomarker for systemic health—ranging from neurological degeneration and respiratory pathologies to mental health disorders—the global medical community is racing to harness algorithms that can perceive what the human ear never could.
A Legacy of Interdisciplinary Innovation
To understand the magnitude of today’s technological leap, one must first appreciate the rich history of the institution driving it. In 1969, Dr. Wilbur James Gould founded The Voice Foundation in New York City. At that time, a truly interdisciplinary approach to the human voice was virtually nonexistent. Medical specialists operated in silos, separated from speech-language pathologists, vocal pedagogues, and performing artists.

Dr. Gould possessed the groundbreaking foresight to break down these institutional barriers, bringing together physicians, scientists, therapists, and performers to share expertise in the holistic care of the professional voice user. This vision materialized in the Foundation’s first Annual Symposium—Care of the Professional Voice—held in 1972, followed by its inaugural Gala (later known as Voices of Summer) in 1973. For more than five decades, the Foundation has successfully built vital bridges between art and science, and between the clinical laboratory and the theatrical stage.
Since 1989, The Voice Foundation has been steered by Dr. Robert Thayer Sataloff, an internationally renowned otolaryngologist who uniquely bridges the worlds of medicine and art as a professional singer and conductor. Having authored more than 1,200 publications, including 79 textbooks, Dr. Sataloff guided the Foundation’s relocation to Philadelphia, where it has expanded its footprint exponentially. Today, the annual Philadelphia Symposium draws hundreds of medical, scientific, academic, and speech-language professionals, alongside elite performing artists from every corner of the globe. Alongside this gathering stands the Journal of Voice, universally recognized as the premier peer-reviewed organ dedicated entirely to voice science and medicine.
Detailed Chronology: Thirty Years of AI Research in Voice Science
While mainstream society only recently awakened to the capabilities of machine learning with the advent of modern large language models, the Journal of Voice has been quietly publishing AI research for over three decades.
The historical timeline of artificial intelligence in voice science began in earnest in 1994—thirty-one years ago—when the journal published a landmark paper by Rihkanen and colleagues titled "Spectral Pattern Recognition of Improved Voice Quality." Published the very same year the World Wide Web was emerging into public consciousness, this study utilized early neural networks to analyze acoustic voice parameters. Voice scientists were exploring artificial intelligence before most working professionals had ever sent an electronic mail.

However, the trajectory of this research remained gradual for many years before hitting a vertical wall of exponential growth. A breakdown of publication volumes across historical blocks tells a dramatic story:
- 1994–2015 (The Foundation Era): 17 papers published over two decades, establishing basic neural network applications and spectral pattern recognition.
- 2016–2019 (The Machine Learning Expansion): 27 papers published as advanced machine learning algorithms began entering clinical workflows.
- 2020–2022 (The Pandemic Catalyst): 16 papers published during a period marked heavily by remote healthcare adaptations and early remote respiratory screening models.
- 2023–2025 (The Exponential Explosion): 102 papers published in just three short years.
In the year 2025 alone, the Journal of Voice published 51 AI-related papers—surpassing the total number of AI papers published in the entire first twenty years of the movement combined. This is not incremental progress; it is an epochal transformation of an entire medical specialty.
Supporting Context and Metrics: Real-World Impact and Global Reach
The significance of these numbers extends far beyond raw publication counts. Academic literature can often languish in obscurity, but the AI research housed within the Journal of Voice is being actively consumed, cited, and deployed by clinicians and engineers worldwide.
Usage statistics underscore this extraordinary global reach. The average AI-focused paper published in the journal has been downloaded nearly 1,000 times—a high threshold for specialized medical literature. More strikingly, the journal’s most-accessed AI paper—a pioneering study examining machine learning applications for COVID-19 detection—has been downloaded over 7,200 times, eclipsing the median for articles in its issue by a factor of eleven.

Similarly, the journal’s most-cited works have achieved massive academic and clinical resonance. A deep learning study authored by Fang and colleagues has amassed 193 citations, while a comprehensive machine learning survey by Hegde has garnered 147 citations. Both papers have been downloaded nearly 5,000 times each and consistently rank among the top-accessed pieces in their respective volumes. These metrics reflect a community of clinicians, researchers, and speech-language pathologists actively seeking out practical, algorithm-driven tools to elevate patient care.
This scholarly ecosystem is propelled by a truly global network of researchers. Key contributors leading the charge include Jérôme René Lechien (6 papers), Dimitar Deliyski (5 papers), Stephanie Zacharias (5 papers), and Ahmed Yousef (4 papers as first author), alongside Paavo Alku, Leonardo Wanderley Lopes, and Maryam Naghibolhosseini (4 papers each). These investigators, working alongside dozens of multidisciplinary labs spanning every inhabited continent, demonstrate that the AI voice revolution is a collaborative, borderless endeavor.
Landmark Research Reshaping the Field
Deep Learning for Voice Pathology Detection
In 2019, Fang and colleagues published "Detection of Pathological Voice Using Cepstrum Vectors: A Deep Learning Approach," which stands as the most-cited AI paper in the journal’s history. By demonstrating that deep neural networks could identify vocal pathologies with remarkable accuracy using acoustic features entirely imperceptible to the human ear, the study opened a new frontier in diagnostic precision.
This work built upon earlier foundational surveys, such as Hegde and colleagues’ comprehensive mapping of machine learning approaches for voice disorders, and Al-Nasheri’s studies on Multidimensional Voice Program parameters. Over the years, the methodology has evolved rapidly. Researchers like Chen and Chen introduced advanced deep neural networks for voice classification, Fujimura pioneered one-dimensional convolutional neural networks (CNNs), and Cho and Choi rigorously compared CNN models for the analysis of complex laryngoscopic images.

Voice as a Digital Biomarker
Perhaps the single most transformative concept emerging from this body of research is the recognition of the voice as a digital biomarker—a non-invasive, continuous window into systemic human health.
Groundbreaking work published in the journal has explored the use of automated voice analysis to detect and monitor neurodegenerative conditions such as Parkinson’s disease. For instance, a 2022 paper by Hemmerling and Wójcik-Pędziwiatr focused on predicting Parkinson’s disease severity directly from voice signals, garnering dozens of citations in a remarkably short timeframe.
Beyond neurology, recent systematic reviews have investigated voice quality as a digital biomarker for psychiatric conditions, including major depressive disorder and bipolar disorder. The clinical implications are profound: routine, automated voice screening could flag early indicators of mental health deterioration, enabling timely medical intervention and easing the strain on overstretched healthcare systems. Additional research deploying machine learning to assess respiratory health from acoustic signatures further highlights the voice’s role as a vital physiological sentinel.
AI Chatbots in Clinical Decision-Making
The year 2025 marked a new phase of integration with the publication of papers examining conversational AI and large language models in clinical workflows. A study by Dronkers and colleagues titled "Evaluating the Potential of AI Chatbots in Treatment Decision-making for Acquired Bilateral Vocal Fold Paralysis" quickly accumulated 18 citations within its first year of publication. This paper sparked lively scholarly discourse, including multiple letters to the editor and follow-up studies evaluating the diagnostic accuracy of platforms like ChatGPT-4o in analyzing complex laryngeal imagery. The field is actively grappling with the ethical, practical, and clinical realities of conversational artificial intelligence in real time.

Expert Perspectives on AI in Voice Science
To contextualize these monumental shifts, leading minds in the discipline offer critical insights regarding the integration of computational tools into medical science.
Dr. Mark Berardi: AI as a Tool for Complexity
Dr. Mark Berardi brings a rare dual perspective rooted in physics and computation to the field of voice science. He views artificial intelligence fundamentally as an instrument for managing complexity.
"Machine learning and AI are tools to help with complexity, and communication is incredibly complex in its neurobiological and physiological processes. So I think the application is warranted," Dr. Berardi notes.
His ongoing research targets voice-based digital biomarkers for aging and depression. While acquiring raw speech and voice signals has become technologically straightforward, the true bottleneck remains the sheer intricacy of human communication—a challenge perfectly suited to machine learning capabilities. Furthermore, Dr. Berardi highlights how generative AI has streamlined his own academic workflow, allowing him to bypass traditional coding bottlenecks by generating and refining bespoke data-processing scripts through natural language prompts.

Intriguingly, Dr. Berardi’s work also extends to the study of human-AI communication itself. As society adapts to "synthetic" communication mediums—such as video conferencing software and automated chatbots—linguistic and vocal patterns are demonstrably shifting. The scientific community may soon need to expand its purview to understand not only how humans speak to one another, but how the human voice adapts when interacting directly with machines.
Dr. Eric Hunter: Embracing Change with Ethical Clarity
Dr. Eric Hunter emphasizes the broad institutional and infrastructural implications of artificial intelligence integration within academia and clinical environments.
"These tools aren’t just novelties," Dr. Hunter observes. "They are rapidly becoming integrated into how we conduct literature reviews, organize data, extract themes from large bodies of text, and even summarize findings."
Large language models and AI-driven platforms are swiftly embedding themselves into standard academic workflows, from manuscript preparation and peer review to collaborative writing and data synthesis. Rather than adopting a stance of technological resistance, Dr. Hunter issues an urgent call to action for the establishment of clear, enforceable guidelines regarding appropriate usage.

To navigate this transition successfully, Dr. Hunter outlines four essential guiding principles for academic integrity:
- Absolute Transparency: Clear disclosure of where and how AI tools are utilized during the research and drafting process.
- Human Accountability: Maintaining strict human oversight to verify all data, citations, and analytical conclusions generated with algorithmic assistance.
- Intellectual Rigor: Ensuring that AI serves to amplify—rather than substitute for—critical scholarly thought and clinical judgment.
- Equitable Access: Fostering an inclusive environment where researchers globally can leverage these tools without prohibitive technological barriers.
"Our field will benefit most," Dr. Hunter concludes, "if we embrace the productivity these tools offer while also building a shared ethical framework for their responsible use."
Future Outlook and the Path Forward
As the Journal of Voice steps into its next decade of publication, it remains unswervingly committed to championing rigorous research that pushes the boundaries of medical science while upholding the highest standards of scientific integrity.
However, the rapid acceleration of artificial intelligence presents distinct challenges that require concerted, collective action from the global community:

- First, practitioners and researchers must engage with AI tools thoughtfully. Algorithms are instruments designed to manage complexity; investigators must understand both their immense potential and their inherent limitations, utilizing them to augment human clinical expertise rather than replace professional judgment.
- Second, the global voice science community must actively participate in drafting and refining shared ethical frameworks. As Dr. Hunter notes, guidelines for ethical AI deployment cannot be dictated by a single institution or journal; they require cross-disciplinary consensus.
- Third, scholars must continue sharing their insights. The 161 AI-related papers housed within the journal represent merely the opening chapter of a much larger story. Every clinician, speech-language pathologist, and educator possesses valuable frontline insights capable of advancing our collective understanding.
An Extraordinary Moment
The human voice has endured for hundreds of thousands of years as our primary instrument of connection, emotional expression, and personal identity. It carries our health, our vulnerabilities, and our very essence in ways that no other biological signal can replicate.
Now, for the first time in human history, we possess computational tools sophisticated enough to genuinely decode that complexity—tools capable of revealing what the voice whispers about our brains, our bodies, and our overall wellbeing. We stand on the cusp of an era where technology can extend the reach of expert clinicians, detect pathological changes long before symptoms manifest, and democratize access to high-quality voice care across the globe.
This is an extraordinary moment in time, and the Journal of Voice, backed by the enduring legacy of The Voice Foundation, proudly stands at its center. The next decade promises even greater discoveries, and the medical and artistic communities alike look forward to the breakthroughs yet to come.
About the Author
Ian DeNolfo is Executive Director of The Voice Foundation, which publishes the Journal of Voice. A graduate of The Juilliard School and The Curtis Institute of Music, he formerly performed as a leading tenor at major opera houses worldwide before transitioning his career to executive leadership at The Voice Foundation.
