TY - JOUR
T1 - Data density scaling for text-to-image models on small dataset
AU - Ye, Senmao
AU - Liu, Fei
AU - Zhang, Dawei
AU - Fan, Hehe
AU - Wei, Wei
AU - Artur, Madal
AU - Wang, Hua
AU - Zheng, Zhonglong
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/10/14
Y1 - 2026/10/14
N2 - This paper demonstrates that increasing data density is more computationally efficient than increasing data quantity. Experiments show that the convergence of the scaling law depends on data density rather than data quantity. The variation of data quantity scaling is nearly linear to data quantity as data density remains far from saturation. In downstream applications, small datasets only cover a small portion of the image space. However, quantity scaling with uncurated data increases the density in the entire image space, which is computationally inefficient. To overcome this problem, density scaling with in-domain (ID) data conserves computational resources by avoiding training on out-of-domain (OOD) data. To distinguish between ID and OOD images, our dual-stage OOD detection method consists of a cluster-based detector that exploits intrinsic manifold structures within the data, and a classification-based detector utilizing fine-grained visual categorization. For constructing ID text-to-image pairs, we assume that the retrieved images exhibit sufficient feature-space density and propose a linear interpolation mechanism to expand associated text features. This integrated OOD detection and interpolation pipeline increases data density with minimal extra computational overhead. Additionally, we apply recurrent affine transformation in our model to construct a hierarchical embedding fusion structure to deal with the frequency gap between time and text embeddings. Extensive experiments demonstrate that small diffusion models are able to synthesize high-quality images at significantly reduced computational cost, achieving FID scores of 6.36 on CUB, 9.52 on Oxford, and 3.89 on COCO. Data, code, and pre-trained models are available at GitHub.
AB - This paper demonstrates that increasing data density is more computationally efficient than increasing data quantity. Experiments show that the convergence of the scaling law depends on data density rather than data quantity. The variation of data quantity scaling is nearly linear to data quantity as data density remains far from saturation. In downstream applications, small datasets only cover a small portion of the image space. However, quantity scaling with uncurated data increases the density in the entire image space, which is computationally inefficient. To overcome this problem, density scaling with in-domain (ID) data conserves computational resources by avoiding training on out-of-domain (OOD) data. To distinguish between ID and OOD images, our dual-stage OOD detection method consists of a cluster-based detector that exploits intrinsic manifold structures within the data, and a classification-based detector utilizing fine-grained visual categorization. For constructing ID text-to-image pairs, we assume that the retrieved images exhibit sufficient feature-space density and propose a linear interpolation mechanism to expand associated text features. This integrated OOD detection and interpolation pipeline increases data density with minimal extra computational overhead. Additionally, we apply recurrent affine transformation in our model to construct a hierarchical embedding fusion structure to deal with the frequency gap between time and text embeddings. Extensive experiments demonstrate that small diffusion models are able to synthesize high-quality images at significantly reduced computational cost, achieving FID scores of 6.36 on CUB, 9.52 on Oxford, and 3.89 on COCO. Data, code, and pre-trained models are available at GitHub.
KW - Diffusion
KW - Out-of-domain detection
KW - Text-to-image
UR - https://www.scopus.com/pages/publications/105042380350
U2 - 10.1016/j.neucom.2026.134293
DO - 10.1016/j.neucom.2026.134293
M3 - Article
AN - SCOPUS:105042380350
SN - 0925-2312
VL - 698
JO - Neurocomputing
JF - Neurocomputing
M1 - 134293
ER -