Amin Karimi Monsefi
Machine Learning Researcher · Apple
I am a Machine Learning Researcher on the Apple MIND team, working on the efficiency of large-scale generative models — both diffusion and autoregressive. My research advances efficient and controllable generative models and the representations that power them — few-step diffusion and flow matching across both continuous (vision) and discrete (language) spaces, self-supervised and vision–language pretraining, and the translation of both into high-impact scientific domains. I completed my Ph.D. in Computer Science at The Ohio State University, advised by Professor Rajiv Ramnath.
News
- Sep 2026 New Joining Apple full time as a Machine Learning Researcher on the MIND team, after completing my Ph.D. at The Ohio State University.
- Sep 2026 Three papers accepted at NeurIPS 2026 — two as first author (Trajectory as the Teacher, DACA-GRPO) and one as co-author (TaxaAdapter).
- 2026 Organizing DiffuLM, the NeurIPS 2026 workshop on Diffusion Language Models: Foundations, Efficiency, and Reasoning — Sydney, 12 December 2026.
- 2027 Serving as a reviewer for ICLR 2027, AAAI 2027, and WACV 2027.
- 2026 Serving as a reviewer for NeurIPS 2026, WACV 2026, and BMVC 2026.
- 2026 Recognized as a Silver Reviewer for ICML 2026.
- Jan 2026 FS-DFM — fast and accurate long-text generation with few-step diffusion language models — accepted at ICLR 2026.
- 2025–26 Received the Graduate Research Award from the OSU Department of Computer Science and Engineering.
- Jul 2025 ISOSNet accepted in Biomedical Optics Express.
- Jun 2025 TaxaDiffusion accepted at ICCV 2025.
Research Interests
Diffusion & Flow Matching in Continuous Space
Controllable and physics-aware generative models for images and video, with few-step samplers that approach or surpass thousand-step teachers.
DiReCT · TaxaDiffusion · TaxaAdapter · KnobGen · Controlla
Discrete Diffusion & Flow Matching for Language
Diffusion language models that close the gap with autoregressive systems while staying parallel and bidirectional — step-aware discrete flow matching, trajectory distillation, and RL with per-step credit assignment.
Self-Supervised & Vision–Language Representation Learning
Pretraining objectives that capture fine-grained structure, underpinning downstream generation, recognition, and segmentation.
Measuring What Generative Models Get Wrong
Metrics and diagnostic protocols for failure modes that visual quality scores miss — how detectable a camouflaged animal really is, and whether a generated storyboard still carries the story's meaning.
Applied Generative Learning for Scientific Domains
Translating these methods to biodiversity, 3D medical imaging, and multimodal spatiotemporal prediction for smart mobility.
Selected Publications
- NeurIPS 2026
Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation - NeurIPS 2026
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models - NeurIPS 2026
TaxaAdapter: Scaling Fine-Grained Species Image Generation to the Tree of Life - ICLR 2026
FS-DFM: Fast and Accurate Long Text Generation with Few-Step Diffusion Language Models - Journal 2025 ISOSNet: A Unified Framework for Cone Photoreceptor Detection and Inner/Outer Segment Length Measurement from AO-OCT B-Scans
- ICCV 2025
TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation - ICLR 2025
Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning - KDD 2024
Masked LoGoNet: Fast and Accurate 3D Image Analysis for Medical Domain - KDD 2023 Novel Physics-Based Machine-Learning Models for Indoor Air Quality Approximations
Show more publications
- CVPR-W 2025
KnobGen: Controlling the Sophistication of Artwork in Sketch-Based Diffusion Models - ICLR-W 2025
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks - Journal 2024 Reducing Manual Labeling Requirements and Improved Retinal Ganglion Cell Identification in 3D AO-OCT Volumes Using Semi-Supervised Learning
- SIGSPATIAL-W 2023 CrashFormer: A Multimodal Architecture to Predict the Risk of Crash
- Journal 2023 Smart and Collaborative Industrial IoT: A Federated Learning and Data Space Approach
- SIGSPATIAL 2022 Will There Be a Construction? Predicting Road Constructions Based on Heterogeneous Spatiotemporal Data
Full list on the publications page and on Google Scholar.
Experience
Machine Learning Researcher — Apple, MIND Team
- Research on the efficiency of large-scale generative models — diffusion and autoregressive.
- ML Research Intern — Apple, MIND Team
- Developed FS-DFM, a step-aware discrete flow-matching framework that matches the quality of 1024-step diffusion baselines in 8 steps (128× speedup), outperforming LLaDA-8B and Dream-7B while being 40× smaller. [ICLR 2026]
- Developed TS-DFM, which replaces blind stochastic jumps in trajectory construction with energy-guided navigation, letting a distilled 8-step student surpass its 1024-step teacher. [NeurIPS 2026]
- Designed DACA-GRPO, a reinforcement-learning framework for diffusion language models using per-step credit assignment and stratified likelihood estimation, improving reasoning on MATH-500, GSM8K, and Sudoku at zero extra inference cost. [NeurIPS 2026]
- Machine Learning Intern — Higharc
- Conducted research on semantic and panoptic segmentation for architectural floor plans.
- Pre-trained a DETR-based model on unlabeled data, addressing the scarcity of labeled examples with self-supervised learning.
- Implemented domain-adaptation approaches to generalize models across datasets with distinct distributions, and strategies to transfer a trained model between domains.
- Senior Machine Learning Engineer — JIBB
- Built computer-vision pipelines for object detection and dynamic content filtering across images and video for a handwriting-capture platform.
- Developed custom CNN architectures to detect content color and remove shadows and reflections.
- Created automated tooling that improved visual clarity in real-time handwriting sessions.
- CTO — BlueBitSoft
- Designed the high-level architecture for pharmacy software solutions, targeting scalability, reliability, and efficiency.
- Aligned technical strategy with business goals across software and domain-expert teams.
- Introduced agile practices and CI/CD pipelines, and led work on performance, security, and regulatory compliance.
- Senior Data Scientist & Back-End Developer — TAPSI
- Developed AI-powered pricing microservices in Python, communicating over RabbitMQ for real-time fare adjustment.
- Designed a GPS anomaly-detection system to prevent fraud and protect rider safety.
- Built data-driven recommendation features (origin, destination, favorite places) using unsupervised learning.
- Created an ETA microservice from live driver GPS traces and published the underlying method.
- Engineered a spatiotemporal forecasting tool to predict high-demand ride areas across urban regions.
Academic Service
Organizing
Reviewing
| Venue | Reviewing | Recognition |
|---|---|---|
| NeurIPS | 2026 | — |
| ICML | 2026 | Silver Reviewer (2026) |
| ICLR | 2025, 2026, 2027 | — |
| AAAI | 2027 | — |
| CVPR | 2025, 2026 | — |
| ECCV | 2026 | — |
| WACV | 2025, 2026, 2027 | — |
| BMVC | 2026 | — |
| ACM SIGKDD | 2024, 2025, 2026 | Outstanding Reviewer (top 10%, 2025 second round) Excellent Reviewer (top 20%, 2025 first round; 2026) |
Awards & Honors
- 2025–26 Graduate Research Award, Department of Computer Science and Engineering, The Ohio State University. Selected by the OSU CSE Department for distinguished research contributions in generative modeling.
- 2026 Silver Reviewer, ICML 2026.
- 2025 Outstanding Reviewer (top 10%) and Excellent Reviewer (top 20%), ACM SIGKDD.
- 2022 Student Travel Award, 30th ACM SIGSPATIAL Conference.
- 2009 Bronze Medal, University of Waterloo Mathematics Olympiad.
Education & Early Research
- Ph.D. in Computer Science — The Ohio State University, Columbus, Ohio
- M.Sc. in Computer Engineering (Software) — Shahid Beheshti University, Tehran
- B.Sc. in Computer Engineering (Hardware) — Shahid Beheshti University, Tehran
