$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Stage art design is a complex multidisciplinary field that integrates scenography, lighting, set design, and spatial storytelling1,2. Historically, the development of stage concepts has relied on human ingenuity expressed through physical models, hand-drawn sketches, mood boards, and iterative brainstorming sessions3,4. This evolution has been profoundly shaped by the tension between artistic vision and the material constraints of the performance space5. As performance requirements grow increasingly sophisticated, the demand for resource-efficient, visually rich, and rapid concept generation has outpaced traditional manual workflows6,7. The emergence of generative artificial intelligence (GenAI) and large language models (LLMs) offers a transformative path to augment the creative process, facilitating design ideation and co-creation in architecture and entertainment8,9.
Traditional scenography often faces a semantic gap where the abstract emotions of a theater script struggle to find a precise visual equivalent without extensive trial and error2,10. Current AI tools in the creative industries allow for the flexible search and recombination of reference photos, yet many lack the structural rigor required for professional stage production11,12. Furthermore, the subjective nature of aesthetic evaluation often makes it difficult to align machine-generated outputs with specific narrative intent13. This shift from experience-driven manual workflows to data-driven human-AI collaboration is becoming a defining characteristic of 21st-century design practice14,15. The integration of these computational tools necessitates a new form of "practice-based" research that bridges the gap between digital tool development and live performance16. Recent studies emphasize that AI systems can facilitate brainstorming and broad exploration of alternatives, yet the transition from 2D ideation to 3D optimized realization remains technically challenging9.
The generative stage art design (GSAD) framework addressed these shortcomings by establishing a repeatable AI-assisted workflow17,18. The rationale for this framework stems from the need to automate design components while preserving the artistic integrity of the original narrative. By integrating multimodal AI systems, designers can quickly iterate both visually and verbally, ensuring that lighting and texture refinements are semantically aligned with the dramatic content19. A critical component of this innovation is the use of the intelligent elephant clan optimization (IECO) algorithm, which draws inspiration from the social behavior of elephant populations to navigate complex spatial constraints in stage layout and lighting placement20,21. This class of metaheuristic algorithms has proven particularly effective in solving multi-objective design problems where traditional gradient-based methods fail.
In the broader context of literature, AI-powered computer-aided design (CAD) systems have already demonstrated efficiency gains of up to 20% in architectural workflows22,23. Specifically, AI-driven generative design allows for the exploration of non-standard spatial topologies that would be impossible to conceive manually24. While recent studies over the past two years have successfully utilized conditional GANs (CGANs) to automate 2D planar floorplan generation and architectural spatial layouts, extending these 2D generative principles to the complex, three-dimensional spatial requirements of stage lighting and set design remains critically under-explored. However, traditional deep learning models, such as convolutional neural networks (CNNs), often struggle with the long-range dependencies required to understand complex 3D stage arrangements25,26. The GSAD framework overcomes these limitations by utilizing vision transformers (ViT) to extract global visual features and bidirectional encoder representations from transformers (BERT) encoders to capture narrative semantics27,28. As transformers capture the global context of a design scene more effectively than local filters, these models provide a more holistic representation of spatial relationships29. This dual-stream feature extraction ensures that the resulting stage designs are not only aesthetically consistent but also narratively relevant.
The advantages of the GSAD approach over alternative techniques, such as the deep convolutional embedding attention mechanism (DL-CBAM), are multifaceted30,31. While previous models improved recognition accuracy on motion detection datasets, these previous models often required higher computational overhead and lacked the spatial optimization capabilities found in GSAD. By contrast, GSAD achieves a 98.88% accuracy rate while maintaining a parameter count of 1.08 M, striking an optimal balance between model complexity and predictive capability32,33. Such efficiency is critical for real-time creative tools, as demonstrated by the success of mobile-scale architectures in specialized design tasks34. This makes the framework highly appropriate for professional theater designers, digital art creators, and architectural engineering educators who require high-fidelity, reproducible visualizations35.
Furthermore, the integration of GANs for texture and lighting refinement addresses the limitations of standard diffusion models, which can sometimes produce "stylistically flat" outputs36,37. By utilizing a style-based generator, the framework can synthesize high-resolution textures that exhibit both variety and artistic depth38. By employing adversarial training, the GSAD framework ensures that artistic details and visual realism are enhanced to professional standards39. This approach also mitigates ethical concerns regarding authorship and bias by providing a controllable, "designer-in-the-loop" collaboration platform40,41. The ultimate goal is not to replace human intuition but to provide a collaborative partner that extends the creative agency of the artist42. As the field moves toward more immersive and digitally integrated performances, the need for robust, AI-native design protocols will only increase43. The following protocol provides the technical roadmap for implementing the GSAD framework, ensuring that modern stage art design remains both imaginative and reproducible.