By Ian DeNolfo
Executive Director, The Voice Foundation
Executive Overview
We stand today at a profound inflection point in the history of voice science—a definitive juncture where artificial intelligence (AI) has ceased to be a futuristic novelty and has instead become a foundational engine of research, data analysis, and patient care.
For over half a century, The Voice Foundation has navigated the space between art and science, bridging the gap between the clinical examination room and the performance stage. Yet, the current paradigm shift driven by artificial intelligence transcends incremental progress; it represents an exponential transformation. From pioneering neural network experiments in the mid-1990s to the explosive surge of deep learning models in the current decade, AI is fundamentally altering how we decode the human voice.
No longer viewed merely as a medium for artistic expression or verbal communication, the human voice is increasingly recognized by modern science as a rich digital biomarker—a non-invasive, highly sensitive window into systemic health, neurological integrity, and mental well-being. As the Journal of Voice—the premier peer-reviewed publication dedicated to this discipline—records record-shattering volumes of AI-focused research, the scientific community must actively embrace these technological capabilities while forging a rigorous ethical framework to guide their deployment.
A Legacy of Innovation: From Gould and Sataloff to the Digital Age
To understand the magnitude of today’s technological revolution, one must appreciate the interdisciplinary bedrock upon which modern voice science was built.

In 1969, Dr. Wilbur James Gould founded The Voice Foundation in New York City. At a time when interdisciplinary care for the human voice was virtually non-existent, Dr. Gould exhibited visionary foresight by bringing together physicians, voice scientists, speech-language pathologists, performing artists, and vocal pedagogues. His goal was simple yet revolutionary: to foster collaborative expertise dedicated to the preservation and care of the professional voice user.
The Foundation inaugurated its first Annual Symposium—Care of the Professional Voice—in 1972, followed by its inaugural Gala (later celebrated as Voices of Summer) in 1973. For more than five decades, these gatherings have served as the intellectual nexus for voice care professionals globally.
Since 1989, the Foundation has operated under the stewardship of Dr. Robert Thayer Sataloff, an internationally renowned otolaryngologist, professional singer, and conductor. Having authored more than 1,200 publications, including 79 textbooks, Dr. Sataloff spearheaded the relocation of the Foundation to Philadelphia, embedding its operations within a vibrant academic ecosystem. Today, the annual Philadelphia symposium draws hundreds of medical, scientific, and artistic professionals from every corner of the globe, while the Journal of Voice maintains its position as the definitive scholarly authority in voice science and medicine.
Detailed Chronology: Thirty Years of AI Research in Voice Science
While public consciousness regarding artificial intelligence—sparked by generative chat interfaces and large language models—is a recent phenomenon, the intersection of AI and voice science has a remarkably deep pedigree within the pages of the Journal of Voice.
The journey began thirty-one years ago, in 1994, when the journal published a seminal paper by Rihkanen and colleagues titled "Spectral Pattern Recognition of Improved Voice Quality." Utilizing rudimentary neural networks to analyze acoustic signals, that research was published the exact same year the World Wide Web first emerged into public view. Voice scientists were exploring artificial neural networks before the vast majority of professionals had ever sent an electronic mail message.

The Exponential Timeline of Publication
For decades, AI research in voice science progressed steadily, if quietly. However, the timeline of recent publications reveals an unprecedented upward trajectory:
- 1994–2015: 17 papers published over two decades, establishing foundational acoustic analysis and early machine learning models.
- 2016–2019: 27 papers published as deep learning architectures began to mature and find clinical utility.
- 2020–2022: 16 papers published amidst global disruptions, laying groundwork for remote diagnostics and tele-health applications.
- 2023–2025: A staggering 102 papers published in just three years, accounting for 63 percent of the journal’s entire historical AI catalog.
The year 2025 alone witnessed the publication of 51 AI-related papers—surpassing the total output of the journal’s first twenty years combined. This is not gradual academic evolution; it is an exponential explosion of data, capability, and clinical translation.
Supporting Context & Metrics: Beyond Citations to Real-World Impact
Academic output is frequently measured in citation counts, but the true metric of success in clinical science is utility: are these studies being read, downloaded, and applied by practitioners in the field?
Usage statistics from the Journal of Voice demonstrate extraordinary global engagement. The average AI-focused paper published in the journal has been downloaded nearly 1,000 times. Outliers achieve staggering reach: a study exploring machine learning applications for COVID-19 detection has amassed over 7,200 downloads—eleven times the median readership for articles within its issue.
Similarly, the journal’s top-cited papers have penetrated deep into clinical workflows. A 2019 deep learning study by Fang and colleagues on pathological voice detection has garnered 193 citations, while Hegde’s comprehensive machine learning survey has accumulated 147 citations, with both papers approaching nearly 5,000 downloads each. These figures reflect a thirsty global community of clinicians, researchers, and speech-language pathologists actively seeking practical, technology-driven tools to elevate patient outcomes.

Landmark Research: Reshaping Clinical Diagnostics
Deep Learning for Voice Pathology Detection
The foundational work of Fang and colleagues demonstrated conclusively that deep neural networks could detect voice pathologies with unprecedented accuracy—often relying on acoustic features and cepstrum vectors entirely imperceptible to the human ear.
This breakthrough paved the way for systematic reviews like Hegde’s survey, alongside foundational contributions from Al-Nasheri et al. (2017) on Multidimensional Voice Program parameters. As the decade progressed, research rapidly evolved: Chen and Chen (2022) advanced deep neural networks for voice classification, Fujimura pioneered one-dimensional convolutional neural networks, and Cho and Choi compared advanced CNN models for the automated evaluation of laryngoscopic imagery.
The Voice as a Digital Biomarker
Perhaps the most paradigm-shifting concept to emerge from recent literature is the conceptualization of the human voice as a digital biomarker—a non-invasive, quantifiable window into systemic pathology.
Significant breakthroughs have been made in utilizing automated voice analysis to detect and monitor neurodegenerative conditions. Hemmerling and Wójcik-Pędziwiatr’s 2022 paper on predicting the clinical severity of Parkinson’s disease strictly from voice signals has secured 29 citations in a remarkably short window, serving as a stepping stone toward automated, point-of-care neurological screening.
Beyond neurology, systematic reviews now explore voice quality as a digital sentinel for psychiatric conditions, including major depressive disorder and bipolar disorder. The clinical implications are profound: routine acoustic monitoring could theoretically flag early shifts in mental health long before traditional clinical presentations occur, alleviating the immense burden currently placed on overstretched mental healthcare systems. Furthermore, machine learning models trained on vocal characteristics during respiratory illnesses (such as COVID-19) highlight the voice’s potential as an early warning system for pulmonary health.

AI Chatbots in Clinical Decision-Making
The integration of conversational AI into medical environments has arrived with remarkable speed. In 2025, the Journal of Voice published groundbreaking research by Dronkers and colleagues titled "Evaluating the Potential of AI Chatbots in Treatment Decision-making for Acquired Bilateral Vocal Fold Paralysis." Accumulating 18 citations almost immediately upon publication, the paper sparked vibrant scholarly discourse, including responses examining the diagnostic accuracy of platforms like ChatGPT-4o in analyzing complex laryngeal images. The medical community is actively grappling with the realities of conversational AI operating within clinical parameters in real time.
Official Statements & Expert Perspectives
To navigate this rapidly shifting landscape, the insights of field-leading pioneers are essential. Two prominent voices within the Journal of Voice community offer distinct, highly complementary perspectives on the integration of artificial intelligence.
Dr. Mark Berardi: AI as a Tool for Complexity
Drawing from a robust background in physics and computation, Dr. Mark Berardi views artificial intelligence not as a replacement for human thought, but as an indispensable instrument for managing overwhelming biological complexity.
"Machine learning and AI are tools to help with complexity, and communication is incredibly complex in its neurobiological and physiological processes. So I think the application is warranted," Dr. Berardi notes.
His research focuses on extracting health insights from aging and depressed voices. While acquiring speech signals has become remarkably simple, the true hurdle remains the intricate nature of the human communication apparatus—precisely the domain where machine learning excels.

Regarding generative AI, Dr. Berardi offers a pragmatic assessment. While it has not yet single-handedly rewritten fundamental research paradigms, it serves as a powerful accelerator against academic bottlenecks. "I can now quickly create bespoke code and edit it with natural language prompts," he explains.
Intriguingly, Dr. Berardi’s current work investigates human-AI communication itself—analyzing how human linguistic patterns shift when interacting with synthetic chatbots or engaging in remote video interfaces like Zoom. As human-machine interaction deepens, voice science must expand its scope to understand not just the isolated human voice, but the acoustic adaptations of the human voice in constant dialogue with algorithms.
Dr. Eric Hunter: Embracing Change with Ethical Clarity
Dr. Eric Hunter focuses heavily on the institutional and operational realities of academic publishing and clinical workflows in the age of AI.
"These tools aren’t just novelties," Dr. Hunter observes. "They are rapidly becoming integrated into how we conduct literature reviews, organize data, extract themes from large bodies of text, and even summarize findings."
Acknowledging that large language models are permanently embedded within academic infrastructure (from Microsoft Office integrations to automated manuscript screening), Dr. Hunter issues a vital call to action. Rather than resisting technological adoption, the academic community must establish clear, transparent usage guidelines.

Dr. Hunter outlines four foundational pillars necessary for responsible academic integration:
- Absolute Transparency: Full disclosure regarding the use of AI tools in drafting, data analysis, or manuscript preparation.
- Human Accountability: Maintaining strict human oversight; authors and reviewers remain entirely responsible for the factual accuracy and integrity of their work.
- Data Privacy & Security: Ensuring patient-derived voice data is protected against unauthorized exploitation or privacy breaches.
- Equitable Access: Ensuring that the benefits of AI-driven voice diagnostics are distributed equitably across diverse global populations and socioeconomic tiers.
"Our field will benefit most," Dr. Hunter concludes, "if we embrace the productivity these tools offer while also building a shared ethical framework for their responsible use."
A Global Research Community
The transformation of voice science is driven by a vibrant, interconnected global network of researchers spanning every continent. The foundational and advanced literature housed within the Journal of Voice is propelled by dedicated scholars, including prolific authors such as Jérôme René Lechien, Dimitar Deliyski, Stephanie Zacharias, Ahmed Yousef, Paavo Alku, Leonardo Wanderley Lopes, and Maryam Naghibolhosseini.
This is not the siloed endeavor of isolated laboratories. It is a unified global enterprise leveraging machine learning to decode the complexities of human vocalization, ensuring that advancements achieved in Philadelphia, Europe, Asia, and the Americas benefit patients worldwide.
Future Outlook: The Path Forward
As the Journal of Voice looks toward the next decade, its commitment to rigorous, peer-reviewed scientific inquiry remains unwavering. However, navigating this new frontier requires collective action across three distinct fronts:

- Thoughtful Engagement with Technology: Practitioners and researchers must view AI as an amplifier of human expertise rather than a substitute for clinical judgment. Understanding both the immense capabilities and the algorithmic limitations of these tools is paramount.
- Collaborative Ethical Governance: The development of a shared ethical framework—as championed by leaders like Dr. Hunter—cannot be achieved by a single journal or institution. It requires an active, cross-disciplinary consensus across medicine, engineering, ethics, and law.
- Continuous Knowledge Sharing: The 161 AI papers published to date represent merely the dawn of a new era. Clinicians, scientists, and educators must continue to publish, debate, and refine their findings to drive the field forward.
An Extraordinary Moment
The human voice has served as our primary instrument of connection, emotional expression, and personal identity for hundreds of thousands of years. It carries the weight of our emotional states, the subtle signatures of our neurological health, and the very essence of our humanity in ways no other physiological signal can replicate.
Now, for the first time in human history, we possess computational tools sophisticated enough to genuinely comprehend that complexity—to decode what the voice whispers about our brains, our bodies, and our overall well-being. These tools empower expert clinicians to extend their reach, catch insidious pathologies before they become catastrophic, and democratize access to world-class voice care across the globe.
We stand at an extraordinary moment in time, and The Voice Foundation alongside the Journal of Voice proudly occupies the epicenter of this revolution. The next decade promises discoveries that will forever redefine the boundaries of voice science and medicine.
