AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier – THE VOICE FOUNDATION

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier – THE VOICE FOUNDATION

Nila Kartika Wati
Nila Kartika Wati

By Ian DeNolfo
Executive Director, The Voice Foundation


Executive Overview

We stand today at a profound inflection point in voice science—a defining historical moment where artificial intelligence (AI) is fundamentally reshaping the paradigms of academic research, data analytics, and patient care. Far from being a sudden trend driven by contemporary generative AI platforms, the integration of computational intelligence into voice science boasts a rich, decades-long heritage. For over fifty years, The Voice Foundation and its premier publication, the Journal of Voice, have served as the epicenter of this interdisciplinary evolution, bridging the gap between artistic expression and rigorous medical science.

This article examines the explosive growth of AI-driven research within voice science, contextualizing its historical roots, analyzing real-world clinical impacts, exploring the concept of the voice as a digital biomarker, and outlining the urgent need for a unified ethical framework. As we navigate an era where machine learning decodes acoustic features imperceptible to the human ear, the imperative for our global community is clear: we must embrace computational complexity while preserving uncompromising scientific integrity and human clinical judgment.

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier - THE VOICE FOUNDATION

A Legacy of Interdisciplinary Innovation

To understand the rapid ascent of artificial intelligence in modern voice research, one must first recognize the visionary foundation upon which it is built. In 1969, Dr. Wilbur James Gould founded The Voice Foundation in New York City. At a time when the interdisciplinary care of the human voice was virtually non-existent, Dr. Gould possessed the extraordinary foresight to assemble a diverse coalition of physicians, scientists, speech-language pathologists, performing artists, and vocal pedagogues. Their shared mission was singular: to revolutionize the clinical and scientific understanding of the professional voice user.

The trajectory of the Foundation was swift and impactful. In 1972, the organization hosted its inaugural Annual Symposium: Care of the Professional Voice, which quickly grew into the premier international gathering for voice professionals. This was followed in 1973 by the first Foundation Gala, an event later celebrated as Voices of Summer. For over half a century, these initiatives have successfully bridged the worlds of art and science, clinic and stage.

Since 1989, The Voice Foundation has operated under the visionary leadership of Dr. Robert Thayer Sataloff, an internationally renowned otolaryngologist, professional singer, and conductor. Having authored more than 1,200 publications—including 79 textbooks—Dr. Sataloff guided the Foundation’s relocation to Philadelphia. Under his stewardship, the institution has continually pushed the boundaries of interdisciplinary scientific research and education. Today, the annual Philadelphia symposium draws hundreds of medical doctors, researchers, academics, speech-language pathologists, and performing artists from across the globe, all anchored by the Journal of Voice, the preeminent peer-reviewed journal dedicated exclusively to voice science and medicine.

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier - THE VOICE FOUNDATION

Detailed Chronology: Thirty Years of AI Research in Voice Science

While public consciousness often views artificial intelligence as a phenomenon born in the 2020s, the Journal of Voice has been pioneering AI-integrated scholarship for more than three decades.

To trace this evolution is to witness the acceleration of computational power meeting biological complexity:

  • 1994 (The Genesis): Thirty-one years ago—coinciding with the public emergence of the World Wide Web—the Journal of Voice published a paper by Rihkanen and colleagues titled "Spectral Pattern Recognition of Improved Voice Quality." Utilizing nascent neural networks to analyze voice signatures, this publication proved that the journal was exploring AI long before digital communication was ubiquitous.
  • 1994–2015 (The Foundational Era): During the first two decades of computational exploration, AI research in the field was sparse yet steady, accounting for 17 foundational papers. Researchers laid the mathematical and acoustic groundwork for automated pattern recognition.
  • 2016–2019 (The Machine Learning Expansion): As computing power expanded and machine learning algorithms matured, the field saw 27 papers published, introducing advanced classification models and laying the groundwork for deep learning applications.
  • 2020–2022 (The Pandemic Catalyst): Amid global shifts toward remote diagnostics and telehealth, 16 papers were published, expanding the scope of machine learning into rapid screening and respiratory assessment.
  • 2023–2025 (The Exponential Explosion): The most recent three-year window has witnessed a staggering paradigm shift, with 102 out of 161 total AI-related papers—63 percent—published during this timeframe. In the year 2025 alone, the journal published 51 AI-related papers, surpassing the total output of the first twenty years combined.

This is not merely gradual academic growth; it is an exponential transformation that mirrors the technological revolutions happening across all biomedical disciplines.

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier - THE VOICE FOUNDATION

Supporting Context & Metrics: Beyond Citations to Real-World Impact

Academic output is frequently measured by citation counts, but the true value of the Journal of Voice‘s AI portfolio lies in its tangible utility for clinicians, researchers, and therapists worldwide. Usage statistics underscore the profound real-world reach of these publications:

  • High Download Velocities: The average AI-focused paper in the journal has been downloaded nearly 1,000 times, reflecting an insatiable appetite among healthcare practitioners for computational tools.
  • The COVID-19 Sentinel Study: The single most-accessed AI paper in the journal’s history—focusing on machine learning methodologies for COVID-19 detection—has amassed over 7,200 downloads, representing eleven times the median readership for articles in its issue.
  • Pioneering Citations: The two most-cited papers in the journal’s AI catalog are Fang and colleagues’ 2019 deep learning study (193 citations) and Hegde’s comprehensive machine learning survey (147 citations). Both articles have been downloaded nearly 5,000 times each, serving as foundational reading for professionals seeking practical diagnostic tools.

Landmark Research Reshaping the Field

The evolution of computational voice research can be traced through specific milestones:

  1. Deep Learning for Voice Pathology Detection: Fang et al. (2019) demonstrated that deep neural networks could detect pathological conditions using cepstrum vectors with remarkable accuracy—often uncovering acoustic anomalies imperceptible to the human ear. This was complemented by Hegde’s expansive survey mapping machine learning approaches for automatic voice disorder detection. Subsequent contributions by Al-Nasheri et al. (2017), Chen and Chen (2022), Fujimura, and Cho and Choi have continuously refined convolutional neural network (CNN) models for laryngeal image and voice classification.
  2. Voice as a Digital Biomarker: Perhaps the most revolutionary concept to emerge from recent literature is the classification of the human voice as a non-invasive digital biomarker for systemic health. Seminal work by Hemmerling and Wójcik-Pędziwiatr (2022) successfully utilized voice signals to predict Parkinson’s disease severity, garnering 29 citations in three years. Furthermore, systematic reviews exploring voice quality as a digital biomarker for depression, bipolar disorder, and respiratory decline demonstrate the potential for routine vocal analysis to alleviate overstretched mental health and primary care systems.
  3. Conversational AI in Clinical Practice: In 2025, the journal expanded its scope to evaluate generative AI and chatbots in clinical decision-making. Dronkers and colleagues’ paper on "Evaluating the Potential of AI Chatbots in Treatment Decision-making for Acquired Bilateral Vocal Fold Paralysis" achieved 18 citations almost immediately, sparking active scholarly debates, letters to the editor, and parallel studies assessing models like ChatGPT-4o in analyzing complex laryngeal imaging.

Official Statements: Expert Perspectives on AI in Voice Science

To contextualize these technological leaps, leading minds within The Voice Foundation community offer critical insights into the philosophical, computational, and ethical dimensions of artificial intelligence.

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier - THE VOICE FOUNDATION

Dr. Mark Berardi: AI as a Tool for Complexity

Drawing from a background in physics and computation, Dr. Mark Berardi views artificial intelligence not as a replacement for human intellect, but as an essential engine for managing biological complexity.

"Machine learning and AI are tools to help with complexity, and communication is incredibly complex in its neurobiological and physiological processes. So I think the application is warranted," Dr. Berardi notes.

Focusing his research on voice-based digital biomarkers for aging and depression, Dr. Berardi highlights the accessibility of vocal data: "Acquiring speech and voice signals is relatively easy now, and we have shown potential for extracting meaningful health information." However, he cautions that the true hurdle remains the intrinsic complexity of human communication itself.

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier - THE VOICE FOUNDATION

On generative AI, Dr. Berardi offers a pragmatic evaluation. While noting it has not yet radically altered foundational research theories, its utility in removing operational bottlenecks is undeniable: "I can now quickly create bespoke code and edit it with natural language prompts." Furthermore, his current research investigates human-AI interaction—examining how human linguistic patterns shift when conversing with chatbots or communicating via synthetic environments like Zoom calls. As human-machine interfaces multiply, voice science must expand to understand speech not just between humans, but within human-machine ecosystems.

Dr. Eric Hunter: Embracing Change with Ethical Clarity

Dr. Eric Hunter focuses his perspective on the institutional and academic realities of AI integration, urging the scholarly community to abandon passive resistance in favor of proactive adaptation.

"These tools aren’t just novelties," Dr. Hunter observes. "They are rapidly becoming integrated into how we conduct literature reviews, organize data, extract themes from large bodies of text, and even summarize findings."

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier - THE VOICE FOUNDATION

As large language models (LLMs) become native fixtures within software infrastructures like Microsoft Office and Google Docs, Dr. Hunter issues a vital call to action: rather than barring these platforms, the academic community must establish clear, transparent guidelines governing their use by authors and reviewers. He advocates for a framework grounded in four essential pillars:

  1. Uncompromising Accountability: Authors remain wholly responsible for the accuracy, validity, and integrity of all published content, regardless of computational assistance.
  2. Transparent Methodology: Any utilization of generative AI in data extraction, coding, or manuscript drafting must be explicitly disclosed.
  3. Intellectual Oversight: AI must serve to augment human analytical depth, never to substitute for critical clinical reasoning.
  4. Equitable Access and Validation: Algorithmic bias must be actively audited to ensure that voice-based diagnostic tools perform reliably across diverse demographic populations.

"Our field will benefit most," Dr. Hunter concludes, "if we embrace the productivity these tools offer while also building a shared ethical framework for their responsible use."

A Global Research Community

The momentum behind this digital transformation is powered by an interconnected international network of scholars. Key contributors advancing AI research within the Journal of Voice include Jérôme René Lechien (6 papers), Dimitar Deliyski (5 papers), Stephanie Zacharias (5 papers), Ahmed Yousef (4 first-author papers), Paavo Alku, Leonardo Wanderley Lopes, and Maryam Naghibolhosseini (4 papers each). Spanning every continent, this global collective demonstrates that the pivot toward computational voice science is a synchronized, worldwide movement.

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier - THE VOICE FOUNDATION

Future Outlook and the Path Forward

As the Journal of Voice charts its course through the twenty-first century, it remains steadfast in its commitment to publishing rigorous, boundary-pushing research while upholding the highest standards of scientific integrity. However, this new frontier demands deliberate, collective action across three critical fronts:

  1. Thoughtful Engagement with Complexity: Researchers and clinicians must master AI platforms as instruments for managing data saturation. We must understand their precise capabilities and inherent limitations, ensuring that technology amplifies human clinical expertise rather than diluting it.
  2. The Establishment of a Shared Ethical Framework: Echoing Dr. Hunter’s insights, no single institution can unilaterally govern the ethical deployment of AI in medicine. The voice science community must unite to forge transparent guidelines for peer review, authorship, and algorithmic accountability.
  3. Continuous Knowledge Sharing: The 161 AI-related papers published to date represent merely the dawn of a new era. Clinicians, speech-language pathologists, and biomedical engineers must continue to share their clinical observations and empirical data to refine diagnostic models.

An Extraordinary Moment

For hundreds of thousands of years, the human voice has served as our primary instrument of connection, emotional expression, and personal identity. It encapsulates our neurological health, our psychological state, and our physiological vitality in a manner unmatched by any other biological signal.

Today, for the first time in human history, we possess technological tools sophisticated enough to truly decode that complexity—to uncover what the voice reveals about our brains and bodies before symptoms manifest, and to democratize access to advanced voice care across the globe.

AI and Professional Voice Care: Journal of Voice 30 Years at the Frontier - THE VOICE FOUNDATION

We are living through an extraordinary moment in medical and scientific history, and The Voice Foundation alongside the Journal of Voice stands proudly at its center. The decade ahead promises discoveries that will redefine the boundaries of medicine, art, and technology. We look forward to pioneering that future together.


  • About the Author: Ian DeNolfo is Executive Director of The Voice Foundation, publisher of the Journal of Voice. A graduate of The Juilliard School and The Curtis Institute of Music, he formerly performed as a leading tenor at major opera houses worldwide prior to his leadership transition within The Voice Foundation.
Your Reaction:

Add a Comment