All reported experiments were executed using the hardware and software environment documented in Table 2 and the Table of Materials. The same graphics driver, CUDA, cuDNN, Python, NumPy, PyTorch, and torchvision versions were used throughout preprocessing, training, inference, and evaluation. The complete hardware and software environment is summarized in Table 2 and the Table of Materials. Table 1 summarizes the processed training, validation, and test datasets together with image-level topology statistics. Table 2 summarizes the common configuration used for all comparison methods.
Comparative evaluation results
The comparative evaluation followed the common data partitions, preprocessing operations, training controls, and metric definitions specified in Protocol sections 7–9. Table 3 compares general visual encoders, part-aligned models, graph-based garment representations, cross-domain fashion retrieval networks, topology-and-appearance fusion methods, and e-commerce visual retrieval methods, all evaluated using the same image-query protocol. The domain-specific methods comprise Topology Feature Histogram with Color and Texture Fusion (TFH-CT), Attention-Guided Cascade Network (AGC-Net), Multi-scale and Multi-granularity Feature Learning Network (MMFL-Net), and Fashion Retrieval using Residual Network with Egret Swarm Optimization (FARE-RNET). All evaluation metrics in Table 3 are reported as the mean and sample standard deviation across five independent runs using the same random seeds, training manifest, validation manifest, test manifest, query-gallery assignment, and metric implementation. The paired p-values were calculated separately for SCI, SDI, Recall@5, mAP, and silhouette coefficient by comparing each method with the proposed method under matched random-seed conditions. These results provide the quantitative basis for the subsequent structural visualization, retrieval, clustering, and design-oriented analyses.
Representative protocol outcomes and interpretation
Successful protocol execution is indicated when the image-level component graph reconstructed from the saved parsing and graph files places every active node inside its corresponding garment region and preserves the relation types recorded in the typed adjacency matrix, as shown in Figure 5C. The structure-aware response should remain concentrated on the garment foreground, as shown in Figure 5D. Figure 6E presents an expected retrieval outcome in which the proposed method reduces the influence of the red background and retrieves upper-body garments rather than unrelated red objects. Figure 6B–G also demonstrates that the retrieval ranking is supported by image-level component relations and post-fusion component correspondence. This result reflects the recovery of the coarse garment category, although sleeve length and local topology still vary among the retrieved samples.
Common failure cases include the texture-dominated activation in Figure 5B, the background-color-driven retrieval in the upper row of Figure 6, and the residual background activation in Figure 7. Figure 7 should be interpreted as partial localization because the dominant response follows the enlarged lower-body silhouette and the right garment boundary, while residual activation remains in adjacent background regions. The localization is considered structurally supported only when the topology-difference matrix, geometric-relation-difference matrix, post-fusion component-difference graph, and spatial response identify consistent garment regions. A spatial hotspot without corresponding changes in the matrix or components indicates visual interference rather than structural evidence.
A protocol run is considered valid when every accepted image produces a normalized semantic probability map, a relation matrix with K rows and K columns, appearance and structural feature matrices with K rows and 512 columns, and a finite fused feature vector linked to the correct image identifier. Visual inspection should confirm that the topology follows the garment components, that the foreground response exceeds the background response, and that retrieval results are guided by garment category and structure rather than by background color. Fragmented probability maps, missing component nodes, nonnumeric matrices, diffuse background responses, or color-dominated retrieval indicate unsuccessful execution and require inspection of parsing, graph construction, feature fusion, or checkpoint loading.
Visual analysis of image-level garment component responses
Figure 5 compares the ResNet-50 baseline with the component-aware model using the same garment image. Figure 5B shows the class activation map for the appearance baseline. Figure 5C presents the image-level garment component graph reconstructed from the foreground-restricted probability map, component center matrix, typed adjacency matrix, inactive-node mask, and component label mapping generated in Protocol sections 3 and 4. Figure 5D presents the activation response of the fused representation. No human pose estimator or human joint annotation is used to generate Figure 5C.
The baseline response was concentrated on local texture regions in the lower garment and did not consistently follow the complete garment boundary, as shown in Figure 5B. In Figure 5C, each node corresponds to the center of an active garment component region, while each edge represents a shared foreground boundary or a predefined vertical-order relation. The graph reflects garment-region organization and does not represent human anatomical joints. The structure-aware response covered the principal garment foreground while maintaining lower activation in most background areas, as shown in Figure 5D. These results indicate that the component graph and feature-fusion module reduced texture-dominated activation in the evaluated image.
Clothing structural similarity analysis and design reference results
Figure 6 presents the retrieval ranking along with the structural evidence used to interpret it. The query component graph in Figure 6B records active garment regions and their typed relations, while the matrix in Figure 6C exposes the relation pattern supplied to graph propagation. The baseline results in Figure 6D remain dominated by red appearance cues. The structure-aware results in Figure 6E retain the upper-body garment category across different colors and backgrounds. The component graph of the highest-ranked result in Figure 6F permits direct comparison with the query graph, and the correspondence matrix in Figure 6G reveals whether components sharing the same identifiers remain aligned after structure-guided fusion. The diagonal correspondence should be interpreted jointly with missing nodes and changed graph relations. Strong appearance similarity without graph agreement does not constitute successful structural retrieval.
The retrieved garments retain the coarse category, but differences in sleeve configuration and local component relations indicate incomplete fine-grained correspondence. The proposed method achieved a Recall@5 of 89.2% and an mAP of 86.4%, as reported in Table 3. These quantitative values describe ranking performance, while Figure 6B–G provides intermediate structural evidence supporting the interpretation. The retrieval conclusion is therefore limited to improved resistance to background-color interference and stronger image-level component organization under the tested query.
Structural difference response and design analysis support
The pure difference-response heatmap in Figure 7H shows that the dominant response is concentrated on the enlarged lower-body silhouette and the right garment boundary of the target image. The overlay in Figure 7I shows that the strongest response remains aligned with the main garment foreground, while weaker activation in the adjacent curtain and wall regions remains as residual background interference. The spatial response should be interpreted together with the component graph and matrix evidence presented in Figures 7C–G. A response is treated as a valid structural difference only when the high-response garment region is consistent with the adjacency change, geometric relation change, and post-fusion component-distance pattern.
A structural difference is accepted only when the typed adjacency change in Figure 7E or the geometric relation change in Figure 7F is accompanied by an increased component distance in Figure 7G and a foreground-aligned spatial response in Figure 7H,I. The enlarged lower-body region and the right garment boundary satisfy this cross-stage criterion. Activation over the curtain and wall does not correspond to changes in component nodes, relation patterns, or fused component distances, and is therefore rejected as non-structural residual interference. The result supports image-level difference localization but does not establish variation in sewing patterns or construction parameters.
Feature spatial distribution and structural cluster analysis
Analyzing the distribution of the latent feature space reveals the model's ability to discriminate the structural semantics of clothing. This analysis examines whether the structural prior organizes garments with different topological configurations into distinct feature manifolds. Five representative structural categories from the test set—short-sleeved tops, trousers, skirts, long coats, and dresses—were selected. The high-dimensional feature vectors were projected onto a two-dimensional plane using t-distributed stochastic neighbor embedding (t-SNE). The cohesion of similar samples and separation of dissimilar samples were assessed to characterize the organization of structural semantics in the feature space. The t-SNE distribution of the structure-aware feature space is shown in Figure 8.
The t-SNE projection formed five main feature regions with limited overlap between neighboring categories, as shown in Figure 8A. Representative samples corresponding to the five category regions are shown in Figure 8B. The SCI of 0.81 reflects the compactness of samples sharing the same structural label, and the SDI of 0.71 reflects the separation between samples with different structural labels, as reported in Table 3 and visualized in Figure 8. These indices should be interpreted as measures of feature-space distribution and do not directly establish a false-alarm rate. Short-sleeved tops and skirts occupied different principal regions in the projected feature space, while a small number of samples remained near neighboring category boundaries, as shown in Figure 8A. Trousers and dresses also showed clear separation from neighboring categories. This distribution indicates that the graph prior contributed to category-level structural organization without producing complete cluster separation. The evaluation separates representation quality from retrieval effectiveness. SCI assesses whether garments sharing the same structural label remain compact in the learned feature space, whereas SDI assesses whether garments with different structural labels remain separated. These metrics correspond to the protocol's structural representation objective. Recall@5 measures whether a structurally relevant garment appears within the first five retrieved results, while mAP measures ranking quality across all relevant gallery samples. These metrics correspond to the design-reference retrieval objective. The silhouette coefficient was used only to evaluate clustering performance because SCI and SDI already characterize feature-space organization.
Figure 9 reports the raw five-run results for four representative methods without cross-method normalization. Figure 9A presents SCI and SDI on their original scale, while Figure 9B presents Recall@5 and mAP as percentages. The individual run markers and standard deviation error lines show the variation associated with model training rather than only the aggregate ranking. The appearance-driven baseline achieved an SCI of 0.62, while the proposed method achieved an SCI of 0.81 and an SDI of 0.71 (Table 3 and Figure 9A). The SCI and SDI results indicate stronger within-label compactness and between-label separation under the evaluated setting. The proposed method achieved a Recall@5 of 89.2% and an mAP of 86.4% (Table 3 and Figure 9B). These retrieval metrics assess different ranking properties and are interpreted separately from SCI and SDI. The consistent ordering of methods across the four metrics indicates agreement between representation quality and retrieval performance, but it does not imply that the relative improvement is identical across the four evaluation dimensions.
Ablation experiments and structural module effectiveness analysis
The ablation experiment compares the complete model with variants in which the structural prior or the fusion module is removed. All three configurations use the same reference and target images, enabling direct comparison of background response and localization concentration. The variant without the structure prior removes the image-level garment component graph and its relation constraints, relying only on the convolutional backbone to extract appearance features; the non-fusion variant retains graph inference but removes the feature pyramid, using only high-level semantic features. Visualized difference heatmaps reveal the specific roles of each module in noise suppression and boundary definition; a qualitative comparison of structural difference responses under different module configurations is shown in Figure 10.
The response maps in Figure 10C–E were rendered using a shared response scale and identical overlay settings. The variant without the structure prior produced strong activation in the left background and around unrelated environmental regions, as shown in Figure 10C. The variant without the fusion module reduced part of that off-garment response but still retained a broader spread over the lower garment and the right-side surrounding region, as shown in Figure 10D. The complete model further contracted the dominant response toward the main garment foreground, as shown in Figure 10E. Figure 10F directly visualizes the difference in foreground-restricted responses between the non-fusion variant and the complete model, showing that the complete model suppresses residual lateral spread while preserving the principal garment-aligned response. These results indicate that the structure prior mainly reduced environmental noise and that the fusion module further improved response concentration and structural alignment. Figure 10E represents a relative improvement rather than a complete removal of background activation.
Table 4 reports SCI, SDI, and Recall@5 for the variants without the structure prior, without the fusion module, and with the complete model. Removing the structure prior reduced the SCI from 0.81 to 0.65 and reduced Recall@5 from 89.2% to 74.5%. Removing the fusion module reduced the SDI from 0.71 to 0.64 and reduced Recall@5 from 89.2% to 85.1%, a decrease of 4.1 percentage points. These results indicate that the structure prior had a greater effect on structural consistency and retrieval, while feature fusion improved the alignment between appearance and structural information.
The efficiency analysis compares the computational cost of the complete model with the ResNet-50 baseline. The ResNet-50 baseline required 25.6 million parameters and 4.1 billion floating-point operations. The complete structure-aware model required 29.3 million parameters and 5.4 billion floating-point operations. The average inference time was 15 milliseconds per image for the baseline and 22 milliseconds per image for the complete model, an increase of 7 milliseconds. This result indicates a measurable computational cost that should be considered when deploying the protocol in time-sensitive workflows (Table 5). The structural constraint weight was determined from the validation split before the final test evaluation. Validation Recall@5 was examined at structural constraint weights of 0.1, 0.3, 0.5, 0.8, 1.0, 1.2, 1.5, and 2.0. A weight of 1.0 was fixed for every subsequent training and test run. Figure 11 documents the validation-based selection of a structural constraint weight of 1.0.

Figure 1: Overall workflow of the structure-aware garment image feature learning protocol. The input garment image is processed by an appearance feature branch and a structural feature branch that uses semantic parsing and spatial graph modeling. The two representations are processed by the structure-guided fusion module to generate the final structure-aware feature. Structural consistency loss, structural discrimination loss, and structural relationship loss jointly constrain the feature space. The figure shows how garment appearance and component topology are integrated throughout feature learning. Please click here to view a larger version of this figure.

Figure 2: Construction of an image-level garment component topology prior. (A) The input garment image I. (B) The foreground-restricted visual component map P, in which seven colors denote image-level garment regions. All background pixels are excluded from the component channels. (C) The image-level component relation graph G. Nodes represent visible garment regions, solid edges represent shared foreground boundaries, and dashed relations represent predefined vertical order. The map and graph support image retrieval and visual structure comparison. They do not represent sewing-pattern pieces, seam geometry, dart structures, grading parameters, or cutting instructions. Please click here to view a larger version of this figure.

Figure 3: Structure-guided feature fusion module. The structural feature is transformed by the linear mapping and activation function Ψ to generate a structure weight. The structure weight modulates the appearance feature through element-wise multiplication. The modulated appearance feature and the structural feature are concatenated and projected by Wc, after which the residual structural feature is added to generate the final component feature . The figure shows that structural information regulates appearance-feature weighting before final fusion. Please click here to view a larger version of this figure.

Figure 4: Feature-space optimization under structural consistency constraints. (A) Samples belonging to the same structural class are pulled closer through the structure consistency loss, while their internal component relationships are preserved by the structure relationship loss. (B) Samples from different structural classes are pushed apart by the structure discrimination loss. The figure shows how the three structural objectives jointly increase within-class compactness and between-class separation. Please click here to view a larger version of this figure.

Figure 5: Representative garment feature responses and component-graph visualization. (A) The input garment image. (B) The class activation response of the ResNet-50 appearance baseline. (C) The image-level garment component graph was reconstructed from the foreground-restricted component map, component center matrix, typed adjacency matrix, inactive-node mask, and component label mapping, which were saved during Protocol sections 3 and 4. Nodes denote active garment regions, solid edges denote shared foreground boundaries, and dashed edges denote predefined vertical-order relations. (D) The activation response of the fused component-aware representation. The comparison shows the difference between texture-dominated activation and garment-region-guided activation. Please click here to view a larger version of this figure.

Figure 6: Retrieval outcomes and intermediate structural evidence under red-background interference. (A) The query image. (B) The query component graph is generated from the saved component centers and the typed adjacency matrix. (C) The query typed adjacency matrix, where the matrix values distinguish inactive relations, shared foreground boundaries, and predefined vertical-order relations. (D) The top-three retrieval results were produced by the appearance baseline. (E) The top-three results were produced by the structure-aware representation. (F) The component graph of the highest-ranked structure-aware result. (G) The post-fusion component correspondence matrix between the query and the highest-ranked result. High diagonal correspondence indicates agreement between components with the same identifiers, while weak or displaced responses indicate missing or mismatched component organization. The retrieval results preserve the coarse upper-body category but do not establish complete agreement in sleeve configuration and local topology. Please click here to view a larger version of this figure.

Figure 7: Structural difference tracing from image-level garment topology to spatial response. (A) The reference image. (B) The target image. (C) The reference component graph. (D) The target component graph. Both graphs are reconstructed from the saved component centers and typed adjacency matrices. (E) The typed adjacency difference matrix. (F) The geometric relation difference is derived from component-center displacement, spatial dispersion, and changes in relation weight. (G) The post-fusion component-difference graph, where node intensity represents the normalized distance between corresponding fused component vectors and edge width represents the relation-weight change. (H) The pure spatial difference-response heatmap was rendered using the same fixed response scale as the overlay. (I) The response overlaid on the target image. A structural difference is accepted only when the adjacency change, geometric relation change, component-distance change, and foreground-aligned heat response remain consistent. Background activation without corresponding evidence of components or relations is treated as residual interference rather than a valid garment-structure difference. Please click here to view a larger version of this figure.

Figure 8: Structure-aware feature-space clustering. (A) The t-SNE projection of short-sleeved tops, trousers, skirts, long coats, and dresses. The five categories form principal feature regions with limited overlap near the boundaries between them. (B) Representative test images assigned to the five garment categories, with the border color corresponding to the category color in (A). The figure shows category-level structural organization while retaining a small number of samples near adjacent category regions. Please click here to view a larger version of this figure.

Figure 9: Raw performance comparison across feature-space structure and retrieval-ranking objectives. (A) SCI and SDI for the appearance-driven baseline, weak-structure model, explicit-structure model, and proposed method. (B) Recall@5 and mAP for the same four methods. Large markers denote the arithmetic mean across five independent runs, error lines denote the sample standard deviation, and small markers denote the run-level results. SCI and SDI are displayed on a fixed scale from zero to one, while Recall@5 and mAP are displayed on a fixed percentage scale from zero to one hundred. No min-max normalization or cross-method rescaling is applied. Please click here to view a larger version of this figure.

Figure 10: Structural difference responses under different module configurations. (A) The reference image. (B) The target image. (C) The response of the variant without the structure prior. (D) The response of the variant without the fusion module. (E) The response of the complete model. (C–E) are rendered using the same normalized response scale, color map, and overlay settings. (F) The foreground-restricted absolute response difference between the non-fusion variant and the complete model. The complete model preserves the principal garment-aligned response while reducing residual spread outside the main structural region. The comparison shows that the structure prior reduces off-garment interference and that the fusion module further sharpens the structural response. Please click here to view a larger version of this figure.

Figure 11: Validation-based selection of the structural constraint weight. Validation Recall@5 is reported at structural constraint weights of 0.1, 0.3, 0.5, 0.8, 1.0, 1.2, 1.5, and 2.0. Each point denotes the arithmetic mean across five validation runs, and each error line denotes the sample standard deviation. The weight was fixed at 1.0 before the final test evaluation. This figure documents the parameter selection procedure and does not constitute an independent architectural contribution. Please click here to view a larger version of this figure.
Table 1: Dataset allocation and image-level topology statistics across the S1 to S5 garment subsets. The table reports the training, validation, test, and total image counts together with the mean semantic-component count, active-node count, typed-edge count, and annotated image-plane landmark count for each structural level. All values are generated from the processed data manifests. Training, validation, and test image counts were obtained from the final subset manifests after structural review, duplicate-group control, and leakage inspection, while topology statistics were calculated from the final graph records using the same component order and relation definitions as those applied during model evaluation. Please click here to download this file.
Table 2: Computational environment, input configuration, augmentation, representation, optimization, loss, and reproducibility settings used throughout the protocol. The table reports the exact operating system, processor, installed memory, graphics processing unit, graphics memory, graphics driver, CUDA, cuDNN, Python, NumPy, PyTorch, and torchvision versions together with the image size, batch construction, optimizer, learning-rate schedule, epoch count, random seeds, representation dimensions, and structural objective weights. Each value is linked to software_versions.txt, configs/protocol.yaml, or the corresponding training record. Please click here to download this file.
Table 3: Quantitative comparison of general visual representation models, graph-based garment models, and domain-specific fashion retrieval methods. The table reports the retrieval focus, query modality, SCI, SDI, Recall@5, mAP, and silhouette coefficient for eleven methods under the same data partitions and evaluation protocol. Every metric is reported as the mean and sample standard deviation across five independent runs. The reported p-values were obtained from two-sided paired t-tests comparing each comparison method with the proposed method, using results produced under the same five random seeds. The proposed method is treated as the reference row. Please click here to download this file.
Table 4: Ablation results for the structure prior and feature-fusion module. The table reports SCI, SDI, and Recall@5 for the variant without the structural prior, the variant without the fusion module, and the complete model. Removing the structure prior yields a larger reduction in SCI and Recall@5, whereas removing the fusion module reduces SDI and response alignment. The complete model records the highest value for all three metrics. Please click here to download this file.
Table 5: Computational efficiency of the ResNet-50 baseline and the complete structure-aware model. The table reports trainable parameters, floating-point operations per image, and average inference time per image under the exact hardware and software environment documented in Table 2 and the Table of Materials. Both models use the same image resolution, inference batch setting, preprocessing operations, and timing procedure. The complete model adds 3.7 million parameters, 1.3 billion floating-point operations, and 7 milliseconds of inference time per image relative to the baseline. Please click here to download this file.