The latest in AI, every dayAI News
Percy Liang

AI Legend · Researcher · Stanford University / Center for Research on Foundation Models

Percy Liang

The man who held a mirror up to AI

August 20, 2026

While the big AI labs competed to build the most impressive model, Percy Liang asked a different question: how do we actually know it works? A Chinese-American researcher, professor at Stanford, and director of the Center for Research on Foundation Models (CRFM), Liang has spent over a decade looking at artificial intelligence from its least glamorous yet most necessary angle: honest, rigorous evaluation.

In 2022, he introduced HELM (Holistic Evaluation of Language Models), the first systematic framework for measuring language models comprehensively. Before HELM, labs picked the benchmarks where their models looked best, and the rest of the world had to take their word for it. HELM changed the rules: it defined 42 real-world application scenarios, seven types of metrics (accuracy, calibration, robustness, fairness, toxicity, bias, and efficiency), and evaluated over 30 models from ten organizations at the same time. It was forced transparency, in a field that doesn't always want to be transparent. Earlier, in 2016, his team created SQuAD (Stanford Question Answering Dataset), the dataset that became the standard for measuring whether a model can read a passage and answer questions about it, driving years of advances in language understanding.

The field has recognized him in multiple ways: he received the IJCAI Computers and Thought Award in 2016, a Sloan Research Fellowship in 2015, a Microsoft Research Faculty Fellowship in 2014, an NSF CAREER Award, and the Presidential Early Career Award for Scientists and Engineers (PECASE) in 2019. He also co-founded Together AI, the platform that brings the best open-source models to production so any company can use them. Liang completed his B.S. and M.Eng. at MIT (2004-2005) and his Ph.D. at UC Berkeley (2011), advised by Michael Jordan and Dan Klein.

To me, Percy Liang's lesson is about intellectual integrity. In an industry where everyone wants to build the next ChatGPT, he prefers to ask the uncomfortable questions: does it actually work? For whom? With what biases? There is no responsible AI without someone willing to measure honestly. Every time you use a language model, part of what makes it trustworthy is the kind of work Liang has been quietly and steadily pushing from Stanford for years.


More figures who shaped AI in AI Legends, or back to the news.