By Ian DeNolfo Executive Director, The Voice Foundation
Executive Overview
Today, we stand at a definitive inflection point in the scientific study of human vocalization. The convergence of artificial intelligence and voice science is not merely augmenting traditional methods; it is fundamentally reshaping how researchers conduct investigations, how data is parsed and interpreted, and how clinicians diagnose and treat vocal pathologies.
For over half a century, The Voice Foundation has occupied the vanguard of this interdisciplinary arena. As the publishing home of the Journal of Voice, the organization has observed—and actively catalyzed—the evolution of voice care from an intuitive art form rooted in empirical observation into a precision science powered by complex computation.
Recent publishing metrics reveal a profound paradigm shift: of the 161 artificial intelligence-focused papers featured in the Journal of Voice since its earliest digital and analog milestones, a staggering 63 percent—102 distinct studies—have been published within the last three years alone. In 2025 by itself, the journal published 51 AI-related papers, surpassing the total publication volume of the field’s first two decades combined.
This is not a story of gradual academic iteration. It is an exponential transformation. As the human voice increasingly proves to be a rich digital biomarker—revealing everything from early-stage neurological decline to systemic mental health conditions and respiratory vulnerabilities—the scientific community is racing to build an ethical, technologically sound framework to harness these revolutionary tools.
A Legacy of Innovation: Bridging Art and Science
To understand the magnitude of today’s technological leap, one must examine the institutional bedrock upon which it rests. In 1969, Dr. Wilbur James Gould founded The Voice Foundation in New York City during an era when comprehensive, interdisciplinary care for the human voice was virtually nonexistent.
Dr. Gould possessed the profound foresight to break down traditional academic and clinical silos. He brought together otolaryngologists, vocal scientists, speech-language pathologists, performing artists, and master voice teachers to share and synthesize their specialized knowledge regarding the professional voice user. This cross-pollination laid the groundwork for the Foundation’s inaugural Annual Symposium—Care of the Professional Voice—in 1972, followed by its first Gala (subsequently christened Voices of Summer) in 1973. For more than five decades, the organization has consistently bridged the chasm between art and science, the clinic and the stage.
Since 1989, this mission has been guided by Dr. Robert Thayer Sataloff, an internationally renowned otolaryngologist, professional singer, and conductor. Having authored more than 1,200 publications, including 79 textbooks, Dr. Sataloff steered the Foundation’s relocation to Philadelphia while expanding its global academic footprint. Today, the annual Philadelphia Symposium draws hundreds of elite medical, scientific, and artistic professionals from across the globe, while the Foundation’s flagship publication, the Journal of Voice, remains the premier peer-reviewed journal dedicated exclusively to voice science and medicine.
Detailed Chronology: 30 Years of AI Research in Voice Science
While contemporary cultural discourse often treats artificial intelligence as a brand-new phenomenon born of modern generative language models, the historical archives of the Journal of Voice tell a radically different story. The journal has championed AI and machine learning research for over three decades—long before the general public understood the architecture of neural networks.
The Pioneers: 1994–2015
In 1994—thirty-one years ago—the Journal of Voice published a seminal paper by Rihkanen and colleagues titled "Spectral Pattern Recognition of Improved Voice Quality." Utilizing primitive neural networks to analyze voice signals, this research arrived during the exact year the World Wide Web was stepping into public consciousness. Voice scientists were exploring machine learning architectures before the vast majority of professionals had ever sent an electronic mail message.
Yet, for more than two decades following that initial spark, publication rates remained modest. Between 1994 and 2015, just 17 AI-focused papers graced the journal’s pages. The technology of the era was constrained by computing power, data storage limits, and feature extraction challenges.
The Acceleration: 2016–2022
As computational power scaled and deep learning architectures matured, interest quickened. Between 2016 and 2019, 27 papers exploring machine learning applications were published. Even amidst the global disruptions of the COVID-19 pandemic from 2020 to 2022, the journal published another 16 papers, laying the technical groundwork for automated acoustic analysis and remote diagnostic screening.
The Exponential Explosion: 2023–2025
Then came the wall of innovation. Between 2023 and 2025, the output exploded to 102 published papers. To contextualize this curve: in 2025 alone, the journal welcomed 51 AI-related manuscripts—outpacing the cumulative output of the first twenty years of the field’s computational history. This trajectory underscores a permanent structural shift in how voice research is conceptualized, executed, and applied.
Supporting Context & Real-World Impact
Academic publication metrics, while illuminating, tell only half the story. The true test of scientific literature lies in its real-world utility—how frequently it is accessed, cited, and integrated into clinical workflows by practicing professionals.
Data from the Journal of Voice demonstrates that these AI-focused papers enjoy extraordinary engagement. On average, an AI article published in the journal is downloaded nearly 1,000 times. Outliers reach astronomical figures: a landmark study investigating machine learning applications for COVID-19 detection has been downloaded over 7,200 times—eleven times the median download rate for articles in its respective issue.
Similarly, the journal’s two most-cited papers—a deep learning study by Fang and colleagues (193 citations) and a comprehensive machine learning survey by Hegde (147 citations)—have each been downloaded nearly 5,000 times. These figures represent active consumption by clinicians, biomedical engineers, and speech-language pathologists seeking practical, empirically validated tools to elevate patient care.
Landmark Research Reshaping the Field
Deep Learning for Voice Pathology Detection
Published in 2019, Fang et al.’s paper, "Detection of Pathological Voice Using Cepstrum Vectors: A Deep Learning Approach," stands as the most-cited AI paper in the journal’s history. The work proved definitively that deep neural networks could isolate vocal pathologies with astounding precision by analyzing acoustic features entirely imperceptible to the human ear.
This foundation was expanded by Hegde and colleagues’ "Survey on Machine Learning Approaches for Automatic Detection of Voice Disorders," which mapped the theoretical and applied landscape of machine learning in vocal health. Concurrently, researchers like Al-Nasheri et al. (2017) pushed multidimensional voice program parameters and correlation functions forward, while newer scholars—such as Chen and Chen, Fujimura, and Cho and Choi—introduced convolutional neural networks (CNNs) for voice classification and automated laryngoscopic image analysis.
Voice as a Digital Biomarker
Perhaps the most transcendent conceptual leap in contemporary voice science is the classification of the human voice as a digital biomarker—a non-invasive, continuous window into systemic health.
The Journal of Voice has published critical research utilizing voice analysis to detect and monitor neurodegenerative conditions like Parkinson’s disease. For instance, Hemmerling and Wójcik-Pędziwiatr’s 2022 study on predicting Parkinson’s severity from acoustic signals has garnered 29 citations in three years, steering the field closer to point-of-care clinical applications.
Beyond neurology, recent systematic reviews highlight the efficacy of vocal quality as a digital biomarker for affective disorders, including depression and bipolar disorder. The clinical implications are immense: routine, automated vocal screening could flag mental health shifts before clinical symptoms fully manifest, drastically reducing the burden on overstretched healthcare systems. Furthermore, machine learning models assessing respiratory health via vocal acoustics have transformed how researchers view the vocal tract—not merely as an organ of speech, but as a sentinel of whole-body wellness.
AI Chatbots in Clinical Practice
The year 2025 marked another milestone with the publication of papers examining conversational AI and large language models (LLMs) in clinical decision-making. Dronkers and colleagues’ study, "Evaluating the Potential of AI Chatbots in Treatment Decision-making for Acquired Bilateral Vocal Fold Paralysis," accumulated 18 citations within months of publication—an extraordinary velocity indicating urgent scholarly interest. This work sparked lively academic discourse, including letters to the editor and investigative responses analyzing ChatGPT-4o’s efficacy in interpreting laryngeal images.
Official Statements and Expert Perspectives
To navigate this rapidly evolving landscape, the voice science community relies on the insights of interdisciplinary leaders who understand both the computational power and the ethical pitfalls of artificial intelligence.
Dr. Mark Berardi: AI as a Tool for Complexity
Dr. Mark Berardi, whose academic roots bridge physics, computation, and voice science, views artificial intelligence primarily as an engine for managing biological complexity.
"Machine learning and AI are tools to help with complexity, and communication is incredibly complex in its neurobiological and physiological processes. So I think the application is warranted," Dr. Berardi explains.
His research centers on extracting health markers related to aging and depression from vocal samples. While acquiring voice data has become frictionless, Dr. Berardi notes that the true bottleneck lies in decoding the intricate mechanics of human communication.
Addressing generative AI, Dr. Berardi offers a pragmatic assessment. While generative models have not entirely rewritten fundamental research methodologies, they have revolutionized workflow efficiencies. "I can now quickly create bespoke code and edit it with natural language prompts," he notes, easing data processing burdens.
Intriguingly, Dr. Berardi’s current work investigates human-machine communication itself—analyzing how human speech patterns adapt when interacting with conversational agents versus living interlocutors, and how synthetic environments like video conferences alter vocal output. "We are already seeing linguistic differences in chatbot interactions," he observes, suggesting that future voice science must encompass human communication with machines.
Dr. Eric Hunter: Embracing Change with Ethical Clarity
Dr. Eric Hunter focuses heavily on the institutional and academic realities of AI integration. He warns against treating advanced computational tools as mere novelties.
"These tools aren’t just novelties," Dr. Hunter observes. "They are rapidly becoming integrated into how we conduct literature reviews, organize data, extract themes from large bodies of text, and even summarize findings."
Dr. Hunter stresses that large language models and automated workflows are permanent fixtures of modern scholarship—embedded deeply within publishing infrastructure, data analysis suites, and collaborative writing platforms. Rather than resisting this technological tide, he issues a clear mandate: the academic community must establish rigorous, transparent guidelines governing appropriate usage for both authors and peer reviewers.
To achieve sustainable integration, Dr. Hunter outlines four essential principles for academic integrity in the age of AI:
Absolute Transparency: Full disclosure of any AI tools utilized during manuscript preparation, data extraction, or coding.
Human Accountability: Maintaining that authors bear total responsibility for the factual accuracy, citations, and intellectual integrity of their work.
Editorial Vigilance: Equipping reviewers and editors with tools and training to detect unverified AI generation or algorithmic bias.
Equitable Access: Ensuring that computational methodologies remain open, reproducible, and accessible to researchers across global institutions.
"Our field will benefit most," Dr. Hunter concludes, "if we embrace the productivity these tools offer while also building a shared ethical framework for their responsible use."
A Global Research Community
The rapid acceleration of AI voice science is propelled by a vast, interconnected international network. Pioneering authors driving this movement within the pages of the Journal of Voice include Jérôme René Lechien (6 papers), Dimitar Deliyski (5 papers), Stephanie Zacharias (5 papers), and Ahmed Yousef (4 papers as first author), alongside Paavo Alku, Leonardo Wanderley Lopes, and Maryam Naghibolhosseini (4 papers each).
Supported by scores of independent laboratories spanning every inhabited continent, these researchers are proving that the digital transformation of voice science is a collaborative global enterprise united by a shared commitment to empirical rigor.
Future Outlook and the Path Forward
As the Journal of Voice steps into the next decade of publication, its editorial mission remains resolute: to champion ground-breaking, methodologically rigorous research while preserving the highest standards of scientific integrity. However, this transition demands proactive, collective engagement from the entire community:
First, researchers and clinicians must engage with AI thoughtfully. These platforms are instruments designed to manage complexity; they must be utilized to amplify human expertise rather than supplant clinical judgment.
Second, the community must co-create universal ethical frameworks. The challenges of algorithmic bias, data privacy, and authorship transparency cannot be solved by a single journal or institution alone.
Third, investigators must continue sharing their findings. The 161 papers published to date represent merely the dawn of a new scientific era. Every clinician, speech-language pathologist, and engineer possesses unique clinical insights capable of advancing our collective understanding.
The human voice has served as our fundamental instrument of connection, emotional expression, and identity for hundreds of thousands of years. It encodes our psychological state, our physiological health, and our very essence in ways no other biological signal can replicate.
Now, for the first time in human history, we possess computational tools sophisticated enough to decode that immense complexity—tools capable of extending the reach of expert clinicians, flagging pathological shifts before symptoms manifest, and democratizing access to premier voice care worldwide.
We are living through an extraordinary moment in scientific history, and The Voice Foundation and its Journal of Voice stand proudly at its epicenter. The next decade promises discoveries we are only beginning to imagine, and we look forward to charting that uncharted territory together.