Theory and Foundations
Masked, uniform, and continuous formulations remain only loosely connected. What is the right unifying formalism, which scaling laws govern training, and which architectural choices actually matter?
NeurIPS 2026 Workshop
About
Language models have been autoregressive by default. Diffusion language models take a different route: they corrupt sequences with noise and learn to reverse it, producing text through iterative refinement rather than one token at a time.
That shift changes what a language model can do: tokens are decoded in parallel, context flows in both directions, and generation becomes far easier to steer and constrain. Over 2025 and 2026 the idea also stopped being theoretical. Inception Labs shipped Mercury, the first commercial-scale diffusion language model; LLaDA and Dream arrived from academic labs; and SEDD, MDLM, FS-DFM, and LaViDa showed that diffusion can match or exceed autoregressive baselines while decoding substantially faster.
Progress has been quick, but the foundations remain fragmented across competing formulations, and the practical questions of efficiency and reasoning are still open. This full-day workshop brings researchers from academia and industry together to consolidate the theory, confront the systems challenges, and examine what diffusion genuinely offers for reasoning.
Masked, uniform, and continuous formulations remain only loosely connected. What is the right unifying formalism, which scaling laws govern training, and which architectural choices actually matter?
Parallel decoding is the central promise, yet practical models still need many denoising steps. How far does that parallelism really extend, and what serving infrastructure does it demand?
Refining a whole solution trace may avoid the cascading errors of left-to-right decoding. Where does that help, where does the lack of sequential structure hurt, and can the diversity gains be measured?
We welcome contributions across the following areas, and related work beyond them.
Timeline
All deadlines are Anywhere on Earth.
Program
Each invited talk is a 30-minute presentation followed by a 5-minute discussion.

Associate Professor, Stanford University
Score-based diffusion and probabilistic inference

MBZUAI Institute of Foundation Models
Foundations and scaling of diffusion language models

Research Director, NVIDIA Research
Scalable generative modeling and generative AI for science

Assistant Professor at Cornell, Co-Founder & Chief Scientist at Inception Labs
Foundations of discrete diffusion: masked, uniform, block diffusion, inference-time scaling, guidance

Research Scientist, Meta Superintelligence Labs
Discrete flow matching and accelerated sampling

Assistant Professor at UCLA, CTO at Inception Labs
Commercial-scale and multimodal diffusion language models
Program
A full day inside the NeurIPS 08:00–17:00 window, with six invited talks, two poster sessions, four oral presentations of accepted papers, and a panel discussion.
Submissions
Submissions may span theory, algorithms, and systems: categorical and masked diffusion; training objectives and their connections to probabilistic inference; sampling, parallel decoding, and efficient inference; reasoning, code generation, and planning; controllable generation and post-training; reinforcement learning and alignment; and hybrid autoregressive and diffusion approaches.
All submissions are non-archival and reviewed double-blind on OpenReview. Work already published at NeurIPS 2026 or another archival venue is not eligible.
Up to 8 pages of main text, with unlimited references and an optional appendix. More complete contributions with experiments or theory.
Every submission receives at least three reviews. We will award a Best Paper and a Best Student Paper prize. Organizers recuse themselves from submissions involving their own institution or active collaborators.
Submission link coming soonCommittee
Listed alphabetically, with complementary expertise across the foundations, efficiency, reasoning, and applications of diffusion language models.

ML Engineering Manager, Apple
LinkedIn ↗
Ph.D. Candidate, HKU
Website ↗
Ph.D. Candidate, Ohio State University
Website ↗
Research Director, NVIDIA Research
Website ↗
Senior Research Scientist, Google DeepMind
Website ↗Assistant Professor, NUS
Website ↗
Research Scientist, Meta
Website ↗
Ph.D. Candidate, Georgia Tech
Website ↗We are actively recruiting reviewers. If you work on diffusion language models or a neighbouring area of generative modeling, we would be glad to have you on the committee.
Sign up as a reviewer