Skip to main navigation Skip to search Skip to main content

Data density scaling for text-to-image models on small dataset

  • Senmao Ye
  • , Fei Liu
  • , Dawei Zhang
  • , Hehe Fan
  • , Wei Wei
  • , Madal Artur
  • , Hua Wang
  • , Zhonglong Zheng
  • Zhejiang Normal University
  • South China University of Technology
  • Zhejiang University
  • Universidade Eduardo Mondlane
  • China-Mozambique “Belt and Road” Joint Laboratory on Smart Agriculture

Research output: Contribution to journalArticlepeer-review

Abstract

This paper demonstrates that increasing data density is more computationally efficient than increasing data quantity. Experiments show that the convergence of the scaling law depends on data density rather than data quantity. The variation of data quantity scaling is nearly linear to data quantity as data density remains far from saturation. In downstream applications, small datasets only cover a small portion of the image space. However, quantity scaling with uncurated data increases the density in the entire image space, which is computationally inefficient. To overcome this problem, density scaling with in-domain (ID) data conserves computational resources by avoiding training on out-of-domain (OOD) data. To distinguish between ID and OOD images, our dual-stage OOD detection method consists of a cluster-based detector that exploits intrinsic manifold structures within the data, and a classification-based detector utilizing fine-grained visual categorization. For constructing ID text-to-image pairs, we assume that the retrieved images exhibit sufficient feature-space density and propose a linear interpolation mechanism to expand associated text features. This integrated OOD detection and interpolation pipeline increases data density with minimal extra computational overhead. Additionally, we apply recurrent affine transformation in our model to construct a hierarchical embedding fusion structure to deal with the frequency gap between time and text embeddings. Extensive experiments demonstrate that small diffusion models are able to synthesize high-quality images at significantly reduced computational cost, achieving FID scores of 6.36 on CUB, 9.52 on Oxford, and 3.89 on COCO. Data, code, and pre-trained models are available at GitHub.

Original languageEnglish
Article number134293
JournalNeurocomputing
Volume698
DOIs
Publication statusPublished - 14 Oct 2026

Keywords

  • Diffusion
  • Out-of-domain detection
  • Text-to-image

Fingerprint

Dive into the research topics of 'Data density scaling for text-to-image models on small dataset'. Together they form a unique fingerprint.

Cite this