TaxaAdapter: Vision Taxonomy Models are Key to Fine-Grained Image Generation over the Tree of Life

Mridul Khurana, Amin Karimi Monsefi, Justin Lee, Medha Sawhney, David Carlyn, Julia Chae, Jianyang Gu, Rajiv Ramnath, Sara Beery, Wei-Lun Chao, Anuj Karpatne, Cheng Zhang

arXiv, 2026

Figure from TaxaAdapter: Vision Taxonomy Models are Key to Fine-Grained Image Generation over the Tree of Life

More than 10M species differ by visual traits so subtle that text-to-image models miss them even when their outputs look photo-realistic. TaxaAdapter injects Vision Taxonomy Model embeddings, such as BioCLIP, into a frozen text-to-image diffusion model, improving species-level fidelity while keeping flexible text control over pose, style and background — and generalising to few-shot and unseen species.

Paper Project page