TaxaAdapter: Vision Taxonomy Models are Key to Fine-Grained Image Generation over the Tree of Life
arXiv, 2026

More than 10M species differ by visual traits so subtle that text-to-image models miss them even when their outputs look photo-realistic. TaxaAdapter injects Vision Taxonomy Model embeddings, such as BioCLIP, into a frozen text-to-image diffusion model, improving species-level fidelity while keeping flexible text control over pose, style and background — and generalising to few-shot and unseen species.
