A fully automated, context-aware pipeline for identity-preserving object insertion from text prompts, fine-tuning Florence-2-Large for spatial grounding and FLUX.1-dev for insertion. Introduces a multimodal data-synthesis framework that improves spatial accuracy, prompt compliance, and visual fidelity without masks, reference images, or manual annotations.
Download PDF