Naman Goyal

Researcher at Google DeepMind and part of the founding team of Gemini Deep Research. I work on long-horizon agents and on the alignment and calibration of large language models.

Before DeepMind I worked on pretraining foundation models at NVIDIA, and interned at Apple (multimodal document understanding) and Adobe Research (adversarially robust metric learning). M.S. in Computer Science from Columbia University, Bachelor’s in Computer Science from IIT Ropar, India, where I graduated Institute Rank 1. Most recently an invited speaker and panelist at ICML 2026 in Seoul.

Papers

Ethan Elasky, Frank Nakasako, Naman Goyal

Can a weaker model judge a stronger one if a critic argues against the answer first? Across five model pairings on verifiable code and logic tasks, debate lifted the judge's reward signal by 7 to 16 points in three pairings and did nothing in two. It helps only when the critic out-classifies the judge and the judge verifies the critique instead of trusting it; a cheap answer, critique, judge loop captures most of the benefit.

Swapnil Parekh, Naman Goyal

Reasoning models keep "thinking" after they have already committed to an answer. ProFIL trains one probe on the frozen base model to detect those post-commitment steps, then filters high-theater rollouts inside GRPO. Across four domains and two model families it cuts reasoning theater by 11 to 100 percent and shortens chains without hurting accuracy.

Naman Goyal, Milan Chaudhari

BinIM binary-searches the adversarial perturbation instead of fixing a step size. On 1,000 ImageNet images it beats FGSM, BIM, and related gradient attacks on three classifiers, driving the true-label probability to about 2e-9.

All papers on Google Scholar

Technical writing

Building a self-correcting code factory where two local LLM agents write, test, and debug Python scripts entirely on Apple Silicon. No cloud, no cost, no data leaving your machine.

All writing

Selected talks

ICML 2026, Trustworthy AI for Good Workshop

Shift Happens: Robustness and Reliability of Multimodal Foundation Models

ICML COEX Grand Ballroom, Seoul · July 2026 · One of the largest ICML 2026 workshops (500+ submissions); keynote by Yoshua Bengio.

ICML 2026, Continual Adaptation at Scale (CATS) Workshop

Shift Happens: Robustness and Reliability of Multimodal Foundation Models

ICML COEX, Seoul · July 2026

Seoul Forum on AI Safety and Security (SFASS), Frontier AI Red-Teaming Workshop

Robustness of Multimodal Foundation Models with Jenny Ni

Seoul · July 2026

Toronto Machine Learning Summit (TMLS) 2026, 10th anniversary

Humans + AI: Collaborative Intelligence for Complex Decision-Making

workshop MaRS Discovery District, Toronto · June 2026 · 90-minute technical workshop; 1,000+ practitioners at the summit.

The AI Conference 2025

The Ascendancy and Challenges of Agentic Large Language Models

Pier 48, San Francisco · September 2025

Videos

Tutorial A 60-minute blitz through gradients and backprop Columbia University, COMS W4732 Computer Vision II · Feb 2022
Talk Humans + AI: Collaborative Intelligence for Complex Decision-Making Toronto Machine Learning Summit (TMLS) 2026, 10th anniversary · Jun 2026
Talk Agentic LLMs in Practice: Tools, Function Calling, and Workflow Design ODSC AI East 2026 · Apr 2026
Talk Taming Non-Determinism: A Framework for Evaluation and Observability in Autonomous Agent Trajectories AIAI Agentic AI Summit Silicon Valley 2026 · Apr 2026
Talk The Ascendancy and Challenges of Agentic Large Language Models The AI Conference 2025 · Sep 2025
Talk Beyond Text Generation: The Rise of Agentic AI and Its Transformative Potential Conf42 Machine Learning 2025 · May 2025

All videos · YouTube channel

All talks

2026

GDM x GDG AI for Science meetup, Google Korea, Seoul Shift Happens: Robustness and Reliability of Multimodal Foundation Modelsevent
Jul 11
ICML 2026, Continual Adaptation at Scale (CATS) Workshop, COEX, Seoul Shift Happens: Robustness and Reliability of Multimodal Foundation ModelsICML pageICML workshopworkshop
Jul 10
ICML 2026, Muslims in ML (MusIML) Workshop, COEX, Seoul Shift Happens: Robustness and Reliability of Multimodal Foundation ModelsICML talk pageworkshop
Jul 9
ICML 2026, Interactive Demos (Google), COEX, Seoul Shift Happens: Robustness and Reliability of Multimodal Foundation Models with Jenny Ni demoGoogle Research
Jul 9
Seoul Forum on AI Safety and Security (SFASS), Frontier AI Red-Teaming Workshop, Seoul Robustness of Multimodal Foundation Models with Jenny Niforumpress release
Jul 8
Toronto Machine Learning Summit (TMLS) 2026, 10th anniversary, MaRS Discovery District, Toronto Humans + AI: Collaborative Intelligence for Complex Decision-Making workshopWatchsummitannouncement
Jun 19
AGI House, Gemini Build Day, Bay Area Building with Gemini 3.5 Flash and the Antigravity harnessevent
May 30
ML Week, Hybrid AI 2026, The Clift Royal Sonesta, San Francisco Taming Non-Determinism: A Framework for Evaluation and Observability in Autonomous Agent Trajectoriessession
May 5
ODSC AI East 2026, virtual Agentic LLMs in Practice: Tools, Function Calling, and Workflow Design workshopWatchevent
Apr 30
AIAI Agentic AI Summit Silicon Valley 2026, San Jose Taming Non-Determinism: A Framework for Evaluation and Observability in Autonomous Agent TrajectoriesWatchsummitspeaker page
Apr 15
DeveloperWeek 2026, AI DevWorld Expo Stage, San Jose Convention Center The Ascendancy and Challenges of Agentic Large Language Modelssession
Feb 18

2025

The AI Conference 2025, Pier 48, San Francisco The Ascendancy and Challenges of Agentic Large Language ModelsWatchslidesspeaker page
Sep 18
Adobe AEP GenAI Seminar Series, Adobe Research World Headquarters, San Jose From Autonomous Reasoning to Collaborative Ecosystems: Architectures for the Next Generation of Enterprise AI Agents
Sep 16
SecurityWeek AI Risk Summit and CISO Forum, The Ritz-Carlton, Half Moon Bay The Ascendancy and Challenges of Agentic Large Language Modelssession
Aug 20
ACL 2025, Google booth, Vienna Gemini Canvas with Heidi ZhangGoogle Research
Jul 29
AI DevSummit 2025, South San Francisco Conference Center The Dual Edge of Multimodal AI: Advancing Accessibility While Navigating Biassession
May 28
Conf42 Machine Learning 2025, virtual Beyond Text Generation: The Rise of Agentic AI and Its Transformative PotentialWatchtalk page
May 8

Judging

Hackathon and competition judging. Open to judging invitations.

GDG Stanford Hackathon, Build with Gemini and Google AI Studio, Stanford Judgeeventluma
May 2026
Google I/O Hackathon, Cerebral Valley x Google DeepMind, SHACK15, San Francisco First-round judge · 100+ builders, $20K+ in prizeseventprojects
May 2026
lablab.ai Agentic Economy on Arc Hackathon, online and onsite Google-track judgehackathon
Apr 2026
Stanford x DeepMind Hackathon, Build with Google AI Studio, Stanford Judge
Apr 2026
YC x Google DeepMind Multimodal Hackathon, San Francisco Judge · Gemini 3.1, Lyria, and Nano Banana 2 early access; listed with the DeepMind team on the event pageevent
Mar 2026

Press

Projects

codex-vision

May 2026

A small Claude Code skill that lets Claude review screenshots, generate UI mocks, and edit images by routing them through OpenAI Codex's vision tools — three modes, one shell command per artifact, gated by a 50-case triggering eval.

Interactive picker for all 228 MacBook Pro M5 build combinations, with live pricing.

Elsewhere