By Ian DeNolfo
Executive Director, The Voice Foundation
Executive Overview
We stand at a definitive inflection point in voice science—a historical juncture where artificial intelligence (AI) has ceased to be an experimental frontier and has instead become a foundational engine of discovery. Across laboratories, clinics, and academic institutions worldwide, AI is fundamentally reshaping how researchers conduct investigations, how data is analyzed, and, most importantly, how patients receive care.
For over five decades, The Voice Foundation and its flagship publication, the Journal of Voice, have served as the premier epicenters for interdisciplinary voice research, bridging the gap between art and science, clinic and stage. Yet, as we examine the trajectory of our field today, we are witnessing a period of unprecedented, exponential transformation. Driven by advanced machine learning architectures, deep neural networks, and large language models (LLMs), voice science is no longer confined to traditional acoustic measurement. Instead, it is pioneering the frontier of digital biomarkers, predictive diagnostics, and AI-assisted clinical decision-making.
This article explores the remarkable convergence of artificial intelligence and voice science, tracing our three-decade history of technological integration, highlighting landmark studies that continue to drive global impact, capturing expert perspectives on the future of research ethics, and mapping the path forward for a truly global, interdisciplinary community.

A Legacy of Innovation: Bridging Art and Science
To understand the magnitude of today’s technological revolution, one must first appreciate the rich legacy upon which it is built. In 1969, Dr. Wilbur James Gould founded The Voice Foundation in New York City at a time when comprehensive, interdisciplinary care for the human voice was virtually nonexistent. With profound foresight, Dr. Gould brought together physicians, vocal scientists, speech-language pathologists, performing artists, and master teachers to share expertise in preserving and restoring the professional voice.
The Foundation hosted its inaugural Annual Symposium—Care of the Professional Voice—in 1972, followed by its first Gala (later christened Voices of Summer) in 1973. Since 1989, the organization has been steered by Dr. Robert Thayer Sataloff, an internationally renowned otolaryngologist, professional singer, and conductor. Under Dr. Sataloff’s visionary leadership, the Foundation relocated its headquarters to Philadelphia and expanded its global footprint, cementing the Journal of Voice as the definitive peer-reviewed organ for voice science and medicine. Today, our annual Symposium in Philadelphia draws hundreds of medical, scientific, academic, and artistic professionals from every corner of the globe.
Detailed Chronology: 30 Years of AI Research in Voice Science
While the broader scientific community often views artificial intelligence as a recent phenomenon born of the mid-2010s deep learning boom, the Journal of Voice has been quietly publishing pioneering AI research for over thirty years.
The Genesis: 1994
In 1994—thirty-one years ago, during the absolute infancy of the World Wide Web—the Journal of Voice published a seminal paper by Rihkanen and colleagues titled "Spectral Pattern Recognition of Improved Voice Quality." This research utilized primitive neural networks to analyze voice signals. To put that in perspective, our field was publishing peer-reviewed artificial intelligence research in voice science before the vast majority of the global population had ever sent an electronic mail.

The Slow Burn to Exponential Growth
For the next two decades, AI research within voice science developed at a steady, incremental pace. However, the publication timeline tells a dramatic story of recent acceleration:
- 1994–2015: 17 papers published
- 2016–2019: 27 papers published
- 2020–2022: 16 papers published
- 2023–2025: 102 papers published
The statistical reality of this growth curve is staggering. Of the 161 total AI-related papers published in the Journal of Voice‘s history, 63 percent (102 papers) have been published in just the last three years. In the single year of 2025, we published 51 AI-related papers—surpassing the total output of the journal’s entire first two decades combined. This is not gradual academic progress; it is an exponential paradigm shift.
Supporting Context & Metrics: Beyond Citations to Real-World Impact
Academic journals are frequently judged by impact factors and citation counts, but the true measure of scientific literature is its real-world utility in clinical practice and ongoing research. The AI portfolio within the Journal of Voice demonstrates extraordinary global reach and engagement.
On average, an AI-focused paper published in our journal is downloaded nearly 1,000 times. However, our most prominent studies far exceed this baseline:

- COVID-19 Detection Study: Our most-accessed AI paper—an investigation into machine learning applications for COVID-19 screening via vocal profiling—has garnered over 7,200 downloads, representing eleven times the median readership for articles in its respective issue.
- Deep Learning Benchmarks: The two most-cited papers in our AI history—Fang and colleagues’ deep learning study (193 citations) and Hegde’s machine learning survey (147 citations)—have each been downloaded nearly 5,000 times. Both ranked among the top three most-accessed articles in their respective issues.
These metrics signify that clinicians, speech-language pathologists, and biomedical researchers are actively seeking and utilizing computational tools to transform patient outcomes.
Landmark Research Reshaping the Field
Deep Learning for Voice Pathology Detection
The cornerstone of modern computational voice analysis is the accurate identification of vocal pathologies from acoustic signals. Fang and colleagues’ 2019 paper, "Detection of Pathological Voice Using Cepstrum Vectors: A Deep Learning Approach," stands as the most-cited AI paper in our journal’s history (193 citations). This work proved conclusively that deep neural networks could detect minute pathological changes in vocal cord function with remarkable accuracy—utilizing acoustic features well beyond the threshold of human auditory perception.
This foundational work has been supported and expanded by comprehensive mapping efforts, such as Hegde and colleagues’ "Survey on Machine Learning Approaches for Automatic Detection of Voice Disorders" (147 citations), alongside pivotal contributions from Al-Nasheri, Chen and Chen, Fujimura, and Cho and Choi, who advanced convolutional neural network (CNN) applications for laryngoscopic imaging and voice classification.
The Voice as a Digital Biomarker
Perhaps the most transformative conceptual leap in modern medicine is the recognition of the human voice as a non-invasive, continuous digital biomarker—a window into systemic neurological, psychological, and respiratory health.

In neurology, recent works—such as Hemmerling and Wójcik-Pędziwiatr’s 2022 study on predicting Parkinson’s disease severity from vocal signals—have rapidly garnered citations and clinical traction, demonstrating how subtle tremors and acoustic instabilities can index disease progression. Beyond neurodegeneration, emerging systematic reviews explore voice quality as a digital biomarker for affective disorders, including depression and bipolar disorder. The clinical implications are profound: routine vocal screenings could eventually allow for early mental health interventions, alleviating the immense burden on overstretched healthcare systems. Furthermore, research assessing respiratory health through acoustic analysis highlights how the voice acts as a sensitive sentinel for systemic physiological stress.
AI Chatbots in Clinical Practice
As conversational AI matures, the Journal of Voice has aggressively engaged with its clinical implications. In 2025 alone, we published a series of papers examining the integration of AI chatbots into clinical decision-making. Dronkers and colleagues’ paper, "Evaluating the Potential of AI Chatbots in Treatment Decision-making for Acquired Bilateral Vocal Fold Paralysis," accumulated 18 citations almost immediately—a rare feat reflecting urgent scholarly demand. Accompanied by rigorous debates and response papers examining platforms like ChatGPT-4o in laryngeal image analysis, our field is actively and transparently grappling with the integration of conversational agents into medical workflows.
Official Statements & Expert Perspectives
To contextualize these technological shifts, we turn to leading voices within our interdisciplinary community, whose insights illuminate both the potential and the responsibilities inherent in this era.
Dr. Mark Berardi: AI as a Tool for Complexity
Dr. Mark Berardi brings a unique perspective rooted in physics and computation. He views artificial intelligence not as a replacement for human intellect, but as an essential instrument for managing biological complexity:

"Machine learning and AI are tools to help with complexity, and communication is incredibly complex in its neurobiological and physiological processes. So I think the application is warranted."
Dr. Berardi’s current research focuses on extracting health insights—such as aging markers and depressive states—from accessible speech signals. While he notes that generative AI has not yet completely rewritten foundational research paradigms, he acknowledges its undeniable utility in solving administrative and technical bottlenecks: "I can now quickly create bespoke code and edit it with natural language prompts."
Furthermore, Dr. Berardi is pioneering research into human-AI communication itself, studying how human linguistic patterns adapt when interacting with synthetic conversationalists or digital environments like Zoom calls. "We are already seeing linguistic differences in chatbot interactions," he notes, suggesting that voice science must soon encompass the study of human speech directed toward machines.
Dr. Eric Hunter: Embracing Change with Ethical Clarity
Dr. Eric Hunter emphasizes the institutional and structural ramifications of widespread AI adoption in academia and clinical practice:

"These tools aren’t just novelties. They are rapidly becoming integrated into how we conduct literature reviews, organize data, extract themes from large bodies of text, and even summarize findings."
Dr. Hunter argues that large language models are permanently embedded within scholarly infrastructure, from manuscript preparation to collaborative writing. Rather than resisting this technological wave, he issues a vital call to action: our community must proactively establish clear ethical guidelines for both authors and reviewers. Dr. Hunter outlines four essential principles for responsible AI integration:
- Transparency: Open disclosure of AI tool usage in data processing and manuscript preparation.
- Accountability: Absolute human oversight and responsibility for generated research outcomes and clinical decisions.
- Data Integrity: Ensuring patient privacy, robust dataset curation, and mitigation of algorithmic bias.
- Editorial Standards: Establishing unified peer-review protocols regarding AI-assisted writing and evaluation.
"Our field will benefit most," Dr. Hunter concludes, "if we embrace the productivity these tools offer while also building a shared ethical framework for their responsible use."
A Global Research Community
The rapid ascent of AI in voice science is propelled by a vast, interconnected network of international researchers. Leading contributors to the Journal of Voice represent institutions spanning every continent. Scholars such as Jérôme René Lechien (6 papers), Dimitar Deliyski (5 papers), Stephanie Zacharias (5 papers), Ahmed Yousef (4 papers as first author), Paavo Alku, Leonardo Wanderley Lopes, and Maryam Naghibolhosseini (4 papers each), alongside dozens of global collaborators, are actively mapping the intersection of machine learning and vocal physiology. This is a unified global enterprise dedicated to decoding the acoustic mysteries of the human condition.

Future Outlook & The Path Forward
As the Journal of Voice enters its next decade, our editorial mission remains steadfast: to publish exceptionally rigorous research that pushes the boundaries of human knowledge while upholding the highest standards of scientific integrity. However, this transition requires collective action across three primary fronts:
- Thoughtful Engagement: We must recognize AI tools for what they are—powerful instruments for managing complexity. We must learn their capabilities, respect their limitations, and use them to amplify human expertise rather than abdicate clinical judgment.
- Ethical Stewardship: Dr. Hunter’s call for a shared ethical framework cannot be answered by a single institution or journal. It demands a concerted, cross-disciplinary consensus among clinicians, engineers, ethicists, and patients.
- Continuous Collaboration: The 161 AI-related papers published in our journal represent merely the foundation. Every clinician, researcher, and educator possesses frontline insights that can advance our collective understanding.
An Extraordinary Moment
For hundreds of thousands of years, the human voice has served as our primordial instrument of connection, emotional expression, and identity. It carries our health, our psychology, and our very essence in ways that no other biological signal can match.
Now, for the first time in human history, we possess technological tools sophisticated enough to truly decode that complexity—to understand what the voice reveals about our brains and bodies, to extend the reach of expert clinicians, to detect pathological shifts long before symptoms become debilitating, and to democratize access to high-quality voice care across the globe.
We are living through an extraordinary moment in scientific history, and the Journal of Voice stands proudly at its center. The next decade promises discoveries we can scarcely yet imagine, and we look forward to sharing every breakthrough with our global community.

About the Author
Ian DeNolfo is Executive Director of The Voice Foundation, which publishes the Journal of Voice. A graduate of The Juilliard School and The Curtis Institute of Music, he formerly performed as a leading tenor at opera houses worldwide before transitioning to arts administration and scientific leadership at The Voice Foundation.
