Abstract
This project explores how structural conditioning can improve controllability in generative image synthesis by comparing text-only diffusion workflows with sketch-guided ones. The experiments were conducted using the Stable Cascade architecture within the ComfyUI node-based environment. In the sketch-guided workflow, simple hand-drawn doodles were processed using Canny edge detection and used alongside text prompts to guide image generation. The goal was to examine how additional structural input influences the diffusion process and whether it improves consistency compared to text prompts alone. To evaluate this, multiple image samples were generated while varying Classifier-Free Guidance (CFG) values and control settings. Both fixed and randomized seeds were tested to observe how these parameters affected prompt adherence, structural stability, and overall composition. Outputs from text-only prompts were compared with those from sketch plus text prompts to understand how multimodal conditioning influences the reliability of generated images. The results show that text-only generation often produces noticeable spatial variation even when seeds remain fixed. In contrast, incorporating sketch-based conditioning provides stronger structural guidance and helps maintain more consistent shapes and layouts. Overall, the findings suggest that combining sketches with text prompts improves structural reliability and user control in diffusion-based image generation workflows.
Advisor
Palmer, Daniel
Department
Computer Science
Recommended Citation
Mac-Iriase, Osen, "Doodle-to-Image: Enhancing Text-Guided Image Generation with Sketch Prompts in ComfyUI" (2026). Senior Independent Study Theses. Paper 13409.
https://openworks.wooster.edu/independentstudy/13409
Disciplines
Artificial Intelligence and Robotics
Keywords
ComfyUI, Stable Diffusion, Image Generation, Stable Cascade
Publication Date
2026
Degree Granted
Bachelor of Arts
Document Type
Senior Independent Study Thesis
© Copyright 2026 Osen Mac-Iriase
