ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Researchers introduced ConVAWG, a retrieval-grounded framework for generating CPS-aligned synthetic multi-turn chat dialogues that model Violence Against Women and Girls (VAWG) scenarios. The framework builds scenarios from persona seeds, UK Office for National Statistics demographic patterns, official crime definitions, and retrieved Domestic Homicide Review cases; converts them into hierarchical event timelines; generates multi-scene role-play dialogues; and applies targeted controls.
ConVAWG is a framework for controlled synthetic dialogue generation in the sensitive domain of Violence Against Women and Girls (VAWG). It addresses the lack of large-scale real conversation datasets due to privacy and legal constraints by generating multi-turn dialogues grounded in official statistics, crime definitions, and real case reviews. The approach models abuse as a relational and temporally unfolding phenomenon, moving beyond sentence-level toxicity detection.
The framework uses retrieval-augmented generation to ground synthetic dialogues in real-world data sources, including demographic patterns and domestic homicide reviews. It structures scenarios into hierarchical event timelines before generating multi-scene role-play dialogues, enabling controlled and contextually rich conversation modeling.
This work highlights a growing need for domain-specific synthetic data generation tools in sensitive areas where real data cannot be easily shared. It may influence how organizations approach data augmentation for training conversational AI in regulated or high-risk domains.
The framework could enable safer and more compliant development of conversational AI systems for social good applications, reducing reliance on scarce or restricted real-world data while maintaining alignment with legal and ethical standards.
Next signals include potential applications in training detection systems for abusive conversations, integration with content moderation pipelines, and extensions to other sensitive domains requiring controlled dialogue generation. Further validation against real-world intervention outcomes would strengthen practical adoption.