The latest in AI, every dayAI News
Kyunghyun Cho

AI Legend · Researcher · NYU Courant Institute / Center for Data Science

Kyunghyun Cho

The man who taught AI to pay attention

August 11, 2026

In 2014, a 28-year-old Korean researcher working as a postdoctoral fellow in Montreal published two papers nobody was expecting, and they changed everything. Kyunghyun Cho was not a well-known name at the time, but the models he created that year live inside every AI application you use today: machine translation, text assistants, language models. Everything starts, somewhere, from what he did that year alongside his collaborators in Yoshua Bengio's lab.

The first paper, presented at EMNLP 2014, introduced the RNN encoder-decoder model and created the Gated Recurrent Unit (GRU), an architecture that allowed a neural network to process variable-length text sequences by compressing them into a vector and reconstructing them in another language. It was the foundation of neural machine translation. But the second paper was the one that lit the fuse: "Neural Machine Translation by Jointly Learning to Align and Translate," written with Dzmitry Bahdanau and Yoshua Bengio, introduced the attention mechanism. Instead of forcing the network to compress an entire sentence into a single vector, it was allowed to look back, searching for which part of the original phrase was relevant to each word being generated. It was an idea that seemed obvious in hindsight, and that nobody had formalized that way before.

Cho is now the Glen de Vries Professor of Health Statistics at NYU's Courant Institute, where he is also a professor of Computer Science and Data Science. His career took him through Aalto University in Finland (where he earned his doctorate in 2014), Bengio's lab in Montreal, and Facebook AI Research, where he was a Research Scientist from 2017 to 2020. He has received the Ho-Am Engineering Prize (2021, one of the most prestigious in Korea), the CIFAR Azrieli Global Scholar Award (2017), and an ICML Best Paper Award (2019). In 2024 he received an ICLR Outstanding Paper Award for work on protein discovery with ML.

What I find fascinating about Cho is that his greatest contribution was conceptual: he wasn't the first to make neural networks bigger, he was the one who made them look better. The attention he formalized in 2014 is the heart of the mechanism that Vaswani and others would turn into the Transformer in 2017, the same one that powers GPT, Claude, Gemini, and everything that came after. Every time a language model understands the context of a long sentence, there is something of Cho's idea working underneath. That is what makes a researcher a legend: it isn't always the most famous name, but the one who opens the door that everyone else needed to walk through.


More figures who shaped AI in AI Legends, or back to the news.