Study design and analytical unit
Each painting or catalog record was treated as one analytical unit. A corpus-based observational design was used, combining metadata review, preliminary computational organization, independent manual coding, and art-historical interpretation. The study period was defined as 1757–1911, and inclusion and exclusion criteria were applied within this timeframe. Historical artworks and their associated catalog metadata were analyzed. When distributions across date intervals of unequal duration were compared, counts were normalized by interval duration, while potential survival and acquisition biases were considered in the interpretation.
Corpus construction and selection
Candidate records were retrieved from the cited museum catalogs, digital archives, and published catalogs. For each candidate record, the institution, title, date or date range, medium, accession or catalog number, stable source URL, retrieval date, and rights statement were retained. Source links were preserved even when the corresponding image file could not legally be redistributed. The reported screening procedure began with 612 candidate images. Of these, 85 were excluded because they fell outside the 1757–1911 period or the export-painting tradition, 58 were excluded because they were modern reproductions, blurred records, or depicted incorrect subjects, 29 were excluded because they lacked an identifiable female figure or adequate image resolution, and 7 were excluded because of insufficient metadata. The final analytic corpus comprised 433 records.
Metadata standardization and image preprocessing
Each painting in the final analytic corpus was assigned a stable QEP identifier ranging from QEP_001 to QEP_433. Associated museum metadata were retained, including institution, accession number, title, date, medium or materials, place of origin, source link, and available catalog description. For visual analysis, the preliminary AI label, adjudicated human category, number of female figures, age group, posture, clothing, spatial setting, associated objects, and social relations were recorded. Each painting was classified into one of seven primary categories: elite/beauty, domestic, servant, laboring, market, performing, or ritual/wedding. For scenes containing multiple female figures, the focal figure was identified according to visual prominence and narrative centrality. When no single female figure was dominant, the primary category was assigned based on the principal depicted activity and the overall scene context. Overlapping categories were resolved by prioritizing depicted activity and scene context over clothing or appearance alone. Ambiguous cases were independently reviewed by both coders and resolved through consensus. Source images were saved in JPEG or PNG format. Only webpage borders, catalog frames, or non-image margins were cropped. Painted content was not reconstructed, recolored, enhanced, or restored.
CLIP-based preliminary labeling
The CLIP ViT-B/32 image encoder was used for preliminary image–text similarity analysis9. Candidate identity terms included elite woman, domestic woman, female servant, working woman, market woman, female performer, bride, and ritual woman. Scene terms included Chinese interior, garden, street market, workshop, tea production, textile work, and ceremonial scene. The complete set of prompt templates, their order, any prompt ensembling, text normalization procedures, decision rules, confidence scores, and ties were retained for each image.
The complete prompt list was applied consistently to all images using the OpenAI CLIP ViT-B/32 model implemented in Python 3.10 with PyTorch 2.0.1 and the OpenAI CLIP package. Inference was performed on a CUDA-enabled GPU with a batch size of 32 using FP32 precision. Each image was resized and center-cropped to 224 × 224 pixels, and the RGB channels were normalized using the preprocessing transformation associated with the ViT-B/32 model. Text prompts were constructed using the identity and scene descriptors defined above and encoded with the CLIP text encoder. Cosine similarity was calculated between normalized image and text embeddings, and the prompt with the highest similarity score was assigned as the preliminary AI label. Similarity scores for all candidate labels were retained for subsequent human review. A fixed random seed of 42 was used, and identical preprocessing procedures, prompt templates, and decision rules were applied to all 433 paintings.
SAM-assisted localization
The Segment Anything Model (SAM) was strictly utilized to localize figures, garments, furniture, tools, ritual objects, and architectural regions for human visual inspection. It must be emphasized that SAM did not participate in the classification pipeline; its masks, crops, and derived features were isolated from CLIP labeling and clustering processes, ensuring that SAM's application provided localization aids without influencing or inflating classification performance.
Manual point prompts or bounding boxes were applied when necessary to refine the localization of visual elements. SAM masks were generated for female figures, garments, furniture, tools, ritual objects, and architectural elements, and each segmentation result was visually inspected for appropriate regional coverage. The resulting masks supported the identification and review of relevant iconographic features but were not used for CLIP labeling or category assignment.
Dimensionality reduction and clustering
CLIP image embeddings were computed, and UMAP with cosine distance was used for two-dimensional visualization10. The representation used for clustering, together with its preprocessing procedure, dimensionality, and distance definition, was recorded.
K-means and hierarchical clustering were applied to the image-feature representations to identify recurrent visual groupings within the corpus11. Candidate K-means solutions for k = 5–9 were evaluated using the silhouette coefficient and Davies–Bouldin index. K-means++ initialization was used with n_init = 10, max_iter = 300, tol = 1 × 10⁻4, and random_state = 42. Hierarchical clustering was performed using agglomerative clustering with average linkage and cosine distance. Candidate solutions were compared using quantitative validation indices, cluster stability, and art-historical interpretability, and the final cluster solution was selected for its most coherent representation of recurrent visual patterns. The resulting cluster assignment for each of the 433 paintings was retained for subsequent comparison with the human-coded categories.
Manual coding, training, and adjudication
Two coders independently reviewed all 433 images using a versioned codebook covering primary category, figure count, age group, posture, clothing, spatial setting, objects, and social relations. Coder-specific records were retained before adjudication. The coders had relevant backgrounds in art history and visual analysis and were trained using the standardized coding manual and representative examples from the painting corpus. Before formal coding, they completed a calibration exercise involving 30 paintings to promote consistent interpretation of the seven identity categories and five visual-cue dimensions.
During formal coding, both coders independently evaluated all eligible paintings while blinded to each other’s coding decisions, preliminary AI labels, and cluster assignments. Inter-coder reliability was assessed using Cohen’s kappa. The observed kappa value for the primary identity category was 0.88, and the categorical visual-cue variables ranged from 0.79 to 0.92, demonstrating strong agreement between coders. Disagreements were resolved through discussion and consensus, and the adjudicated category and visual-cue codes were recorded for subsequent analysis12. Pre-adjudication labels were preserved.
AI–human agreement and performance estimation
The adjudicated human category was used as the reference for descriptive comparison with one preliminary AI label per painting. Overall exact-match accuracy was calculated as 356/433 (82.2%). Category recall was calculated as the number of exact AI–human matches divided by the human-reference support for each category. Category support, exact matches, mismatches, and corresponding denominators were reported explicitly.
Based on the 433 case-level paired AI and human labels, target-class false positives, per-category precision, recall, and F1 scores were computed. A full confusion matrix was constructed to analyze off-diagonal classification patterns, which subsequently guided the mismatch review, error taxonomy documentation, and adjudication rules.
Reproducibility package and materials
The computational workflow and supporting research materials were organized in a version-controlled repository. The reproducibility package was designed to include scripts for preprocessing, CLIP analysis, SAM localization, UMAP, clustering, agreement analysis, and figure generation, together with the manual coding codebook, configuration files, dependency information, case-level derived outputs, and a machine-readable inventory of image sources and rights information. A persistent identifier was assigned to the archived research materials to facilitate reproducibility and reuse.
Microsoft Excel 2021 was used to organize and manage metadata. Computational analyses used CLIP ViT-B/32 for preliminary vision–language labeling, SAM for visual-element localization, UMAP with cosine distance for dimensionality reduction, and K-means and hierarchical clustering for exploratory pattern analysis. Clustering solutions were evaluated using the silhouette coefficient and Davies–Bouldin index, and inter-coder agreement was assessed using Cohen’s kappa. The software environment, model configurations, computational settings, and random seeds used throughout the workflow were documented to support reproducibility.