Recent research has revealed significant challenges in detecting and evaluating hateful intent in multi-turn visual story generation. This study highlights how advancements in text-to-image (T2I) systems, such as Gemini and GPT-Image, have made the creation of hateful narratives, particularly in visual formats like comics and picture books, more accessible and systematic. The authors of the paper, titled "Innocent Panels, Hateful Stories," introduce a new dataset called HatefulStoryPrompts, which includes 330 multi-turn configurations derived from 55 hateful narratives across two languages and three visual styles.
The study meticulously evaluates five leading models against 4,950 story generation attempts, where the best performing model successfully completes 99.0% of the stories. However, existing moderation systems show inadequacies in addressing the group-level meanings of these generated narratives. The evaluation of the HatefulVisualStory dataset, which comprises 969 hateful image sets and 990 benign controls, revealed that dedicated safety models have a recall rate of only 34.9%, while more comprehensive vision-language models achieve 67.5% recall.
To address these limitations, the authors propose proactive and post-generation defenses. An interaction-aware monitoring system demonstrates a 97.3% recall rate for sessions initiated with prompts alone, dropping slightly to 92.6% when users provide the first image. Meanwhile, post-generation evaluations of image groups yield an 80.2% recall rate. The authors conclude that as visual story generation transitions from single images to coherent narratives, safety measures must evolve from simple per-image moderation to a more integrated approach that considers narrative context and inter-image relationships.