Abstract

This project explores how structural conditioning can improve controllability in generative image synthesis by comparing text-only diffusion workflows with sketch-guided ones. The experiments were conducted using the Stable Cascade architecture within the ComfyUI node-based environment. In the sketch-guided workflow, simple hand-drawn doodles were processed using Canny edge detection and used alongside text prompts to guide image generation. The goal was to examine how additional structural input influences the diffusion process and whether it improves consistency compared to text prompts alone. To evaluate this, multiple image samples were generated while varying Classifier-Free Guidance (CFG) values and control settings. Both fixed and randomized seeds were tested to observe how these parameters affected prompt adherence, structural stability, and overall composition. Outputs from text-only prompts were compared with those from sketch plus text prompts to understand how multimodal conditioning influences the reliability of generated images. The results show that text-only generation often produces noticeable spatial variation even when seeds remain fixed. In contrast, incorporating sketch-based conditioning provides stronger structural guidance and helps maintain more consistent shapes and layouts. Overall, the findings suggest that combining sketches with text prompts improves structural reliability and user control in diffusion-based image generation workflows.

Advisor

Palmer, Daniel

Department

Computer Science

Disciplines

Artificial Intelligence and Robotics

Keywords

ComfyUI, Stable Diffusion, Image Generation, Stable Cascade

Publication Date

2026

Degree Granted

Bachelor of Arts

Document Type

Senior Independent Study Thesis

Share

COinS
 

© Copyright 2026 Osen Mac-Iriase