跳至主導覽 跳至搜尋 跳過主要內容

MoEdit: On Learning Quantity Perception for Multi-object Image Editing

  • Macao Polytechnic University
  • Fuzhou University
  • Great Bay University
  • Sichuan University
  • Shanghai Jiao Tong University

研究成果: Conference article同行評審

4 引文 斯高帕斯(Scopus)

摘要

Multi-object images are prevalent in various real-world scenarios, including augmented reality, advertisement design, and medical imaging. Efficient and precise editing of these images is critical for these applications. With the advent of Stable Diffusion (SD), high-quality image generation and editing have entered a new era. However, existing methods often struggle to consider each object both individually and part of the whole image editing, both of which are crucial for ensuring consistent quantity perception, resulting in suboptimal perceptual performance. To address these challenges, we propose MoEdit, an auxiliary-free multi-object image editing framework. MoEdit facilitates high-quality multi-object image editing in terms of style transfer, object reinvention, and background regeneration, while ensuring consistent quantity perception between inputs and outputs, even with a large number of objects. To achieve this, we introduce the Feature Compensation (FeCom) module, which ensures the distinction and separability of each object attribute by minimizing the in-between interlacing. Additionally, we present the Quantity Attention (QTTN) module, which perceives and preserves quantity consistency by effective control in editing, without relying on auxiliary tools. By leveraging the SD model, MoEdit enables customized preservation and modification of specific concepts in inputs with high quality. Experimental results demonstrate that our MoEdit achieves State-Of-The-Art (SOTA) performance in multi-object image editing. Data and codes are available at https://github. com/Tear-kitty/MoEdit.

原文English
頁(從 - 到)2683-2693
頁數11
期刊Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOIs
出版狀態Published - 2025
事件2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, United States
持續時間: 11 6月 202515 6月 2025

指紋

深入研究「MoEdit: On Learning Quantity Perception for Multi-object Image Editing」主題。共同形成了獨特的指紋。

引用此