Jennifer Neville (Microsoft): Evaluating AI Beyond Testing

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Responsible for research at Microsoft and a professor at Purdue, Jennifer Neville describes a journey that has taken her from academic theories to deployed systems. Her current credo revolves around two interconnected axes: measuring what matters to users and scrutinizing the errors that arise outside of benchmarks. An interview with an award-winning expert in learning from structured and interactive data.
From Theory to Real Systems: A Choice Reinforced by Two Sabbaticals
Jennifer Neville joined Microsoft in 2021 while continuing her role as a professor, after twenty years of teaching and mentoring students. Her trajectory alternates between theory and application, a balance honed during two successive sabbaticals. The first, at the Simons Institute in Berkeley, was highly theoretical and facilitated her return to academia. The second led her to Microsoft Research, where observing algorithms confronted with real systems, real users, and real data kept her in the industry. She emphasizes that the difference between academic and industrial research lies in the goals pursued: in a corporate setting, scaling in concrete systems is of a different nature, with a focus on products and business interests altering the questions posed. She advocates for back-and-forth exchanges between academic and industrial worlds, cites the synergies between the two, and believes that working today on AI systems makes the industry the place to be.
Measuring for Users and Examining Failures Outside Benchmarks
Jennifer Neville's priority in current AI is evaluating systems against the concrete needs of users. She highlights surprising failures when testing models beyond traditional references and offers practical advice for interacting with these systems as they are actually used. When results clash with expectations, she recommends a close examination of the data. The lessons she draws from decades of progress in AI fuel this focus on meaningful measurements and the configurations where models fail.
A Path to AI Built Against Initial Choices
As a child, Jennifer Neville wanted to avoid computer science, her father's field. At university, she initially pursued mathematics and then physics, feeling disconnected, and eventually interrupted her studies. Upon returning, she chose cognitive science but found the program too vague and insufficiently mathematical. She believes that advice steering her toward AI, at the intersection of cognitive science, mathematics, and computational thinking, would have saved her time. After a stint in industry, she returned to specialize in computer science to work with data and stumbled upon AI by chance, initially unaware that it fell under computer science.
Research as Practice: From Workshop to the Adrenaline of Discovery
Her entry into research stemmed from a requirement of the honors program in computer science: to conduct a scientific project and define its subject. Interested in data and AI, she focused on knowledge extraction from interconnected web data, under the guidance of a teacher involved in emerging statistical relational learning. The project resulted in a workshop publication and sparked her desire for graduate studies. She describes the dynamic of research work, characterized by loops of unsuccessful trials followed by breakthroughs, as a source of energy: understanding something at the frontier of knowledge has motivated her sustainably.
Fields of Study and Recognitions from a Prolific Career
Jennifer Neville studies machine learning and AI for interactive contexts and structured data, emphasizing the influence of training data on system behaviors and their alignment with user expectations. She has published over 130 papers, amassed more than 10,000 citations, and received a National Science Foundation CAREER Award. She has been featured on the IEEE "10 to Watch in AI" list and received best paper awards at KDD and ICLR. Chad Atalla emphasizes her foresight and the continuity of her guiding thread throughout the evolution of the field. The cited academic funders include DARPA and NSF.
Framework of the Exchange: Evaluation and Challenges Viewed from Microsoft
The discussion is led by Chad Atalla, a principal applied scientist at Microsoft Research, who introduces Jennifer Neville as the head of research at Microsoft Research and a tenured professor at Purdue. Atalla notes that he has been working at Microsoft for just over six years, primarily on AI evaluation, a field that poses many fundamental questions and challenges. He explains that he discovered Neville's contributions because some of her research offers solutions and sheds light on these issues. The exchange more broadly interrogates what we learn when human or artificial trajectories deviate from expectations, a theme in line with Neville's career dedicated to concrete applications of AI.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.