跳至主導覽 跳至搜尋 跳過主要內容

Data density scaling for text-to-image models on small dataset

  • Senmao Ye
  • , Fei Liu
  • , Dawei Zhang
  • , Hehe Fan
  • , Wei Wei
  • , Madal Artur
  • , Hua Wang
  • , Zhonglong Zheng
  • Zhejiang Normal University
  • South China University of Technology
  • Zhejiang University
  • Universidade Eduardo Mondlane
  • China-Mozambique “Belt and Road” Joint Laboratory on Smart Agriculture

研究成果: Article同行評審

摘要

This paper demonstrates that increasing data density is more computationally efficient than increasing data quantity. Experiments show that the convergence of the scaling law depends on data density rather than data quantity. The variation of data quantity scaling is nearly linear to data quantity as data density remains far from saturation. In downstream applications, small datasets only cover a small portion of the image space. However, quantity scaling with uncurated data increases the density in the entire image space, which is computationally inefficient. To overcome this problem, density scaling with in-domain (ID) data conserves computational resources by avoiding training on out-of-domain (OOD) data. To distinguish between ID and OOD images, our dual-stage OOD detection method consists of a cluster-based detector that exploits intrinsic manifold structures within the data, and a classification-based detector utilizing fine-grained visual categorization. For constructing ID text-to-image pairs, we assume that the retrieved images exhibit sufficient feature-space density and propose a linear interpolation mechanism to expand associated text features. This integrated OOD detection and interpolation pipeline increases data density with minimal extra computational overhead. Additionally, we apply recurrent affine transformation in our model to construct a hierarchical embedding fusion structure to deal with the frequency gap between time and text embeddings. Extensive experiments demonstrate that small diffusion models are able to synthesize high-quality images at significantly reduced computational cost, achieving FID scores of 6.36 on CUB, 9.52 on Oxford, and 3.89 on COCO. Data, code, and pre-trained models are available at GitHub.

原文English
文章編號134293
期刊Neurocomputing
698
DOIs
出版狀態Published - 14 10月 2026

指紋

深入研究「Data density scaling for text-to-image models on small dataset」主題。共同形成了獨特的指紋。

引用此