$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
This study does not involve human participants or animal subjects. The analysis is based solely on publicly available legal case materials and curated legal datasets. Therefore, ethics approval and informed consent were not required.
The proposed methodology aims to capture the legal uncertainty surrounding ownership of AI-created content by developing a Bayesian Network based on actual case data. To start with, cases related to AI, including copyright, patent, and trade secrets, would be harvested from an existing corpus and each tagged with actual legal outcome data. Based upon legal precedent and existing case studies, uncertainty parameters that could have an impact on content ownership would include creative human input, AI system autonomy, legality of training data, contractual definitions of ownership, jurisdiction, and transparency, each assigned as nodes within the Bayesian Network with discreet states indicating differing legal and technical realities, and with content ownership determination as the target node. The structure of this network is discovered by integrating legal expertise with data-driven structural discovery to promote both interpretability and legal plausibility. The tables of conditional probabilities are then approximated based on litigation outcomes identified in cases, allowing the network to process uncertain and incomplete data probabilistically. Subsequent to model training, it allows for both inferencing and scenario testing to determine how changes in contributing factors impact inferred ownership in AI-generated content based on this constructed Bayesian Network, followed by testing and sensitivity analysis validation prior to inferring results that facilitate future legal implications for intellectual property in AI applications by developers and in courts. Figure 1 shows the architectural workflow of the proposed work.

Figure 1: Architectural Diagram of the Proposed Framework: Overview of the proposed Bayesian Network–based system for determining ownership of AI-generated content. The architecture illustrates the flow from data acquisition and annotation to probabilistic modeling, inference, and ownership attribution. Please click here to view a larger version of this figure.
Figure 1 illustrates the architectural workflow of the proposed Bayesian Network–based framework for AI-generated content ownership analysis. The workflow includes data collection, feature annotation, Bayesian Network modeling, probabilistic inference, scenario analysis, and validation. Legal and technical variables, such as human involvement, AI autonomy, training data provenance, contractual context, and jurisdiction, are modeled within a Bayesian framework to estimate ownership attribution under uncertain conditions. The framework supports interpretable probabilistic reasoning and scenario-based evaluation for legal decision support.
Data collection
Data on legal cases was collected from various sources to create a dataset focused on ownership and intellectual property lawsuits involving AI-generated content. To start, the general dataset of legal cases sourced in the base paper was filtered to include only litigations concerning copyright, ownership, IP-related disputes, and other IP matters in the training data. Additional legal cases were obtained from publicly available court-opinion databases that provide access to court opinions from millions of cases, as well as a copyright case-tracking repository that tracks both ongoing and concluded legal proceedings worldwide on uses of Artificial Intelligence regarding Intellectual Property. Legal cases were also annotated to highlight key variables, including the legal case outcome, the level of human creative input, the autonomy of the Artificial Intelligence system, the origin of the training data, jurisdiction, and contract parameters, if applicable. This is a multi-source dataset that aggregates sufficient data on Intellectual Property related to the use of AI tools, suitable for probabilistic reasoning in Bayesian network analysis.
Three main sources, including the base AI litigation dataset, a public court-opinion database, and an AI copyright case tracking repository, were used in a methodical collection and screening process to create the dataset for this study. Keywords such as “artificial intelligence,” “AI-generated content,” “copyright,” “intellectual property,” “authorship, “ownership,” and “machine-generated works” were used throughout repository search interfaces and metadata fields to find pertinent cases. Cases involving AI-generated or AI-assisted content that addressed ownership, authorship, or intellectual property concerns with enough depth for variable extraction were considered. Cases that were duplicates, unconnected to AI-generated content, concentrated on other legal areas, or lacked adequate information were not accepted. A two-phase screening procedure was used, with a full-text eligibility assessment coming after the initial title and summary examination. To ensure uniqueness, duplicate records across multiple datasets were identified using case metadata and removed.
The keywords “artificial intelligence,” “AI-generated content,” “machine-generated works,” “copyright,” “authorship,”“ownership,”“training data,” and “intellectual property dispute” were used in the dataset search process. To guarantee the quality and relevance of the dataset, a two-stage screening procedure was used. To find ownership and copyright conflicts pertaining to AI, titles and case summaries were reviewed in the first stage. To determine eligibility for inclusion in the study, a full-text case review was conducted in the second stage. Only cases involving AI-generated or AI-assisted content, with adequate legal and technical details regarding authorship, ownership attribution, contractual context, or training data usage, were included. Excluded were cases with insufficient contextual information, inadequate legal documents, duplicate entries, and cases unrelated to intellectual property conflicts. By comparing metadata such as case title, jurisdiction, filing date, and dispute category, duplicate cases across sources were identified and eliminated. Annotation, Bayesian Network modeling, probabilistic inference, and sensitivity analysis were then performed on the final curated dataset.
The data employed in the study comprises a compilation of legal sources covering the full scope of disputes over AI content ownership rights and intellectual property claims, as shown in Table 2. First, the AI litigation data used in the primary study includes the selective preservation of about 72 legal disputes specifically tied to copyright claims, authorship, authorship by invention, and ownership involving AI outputs. To increase data variability for better representation of emerging disputes in the realm of generative AI, the data includes the acquisition of relevant disputes sourced from the publicly available court opinion database (CAP), a massive publicly available collection of U.S. court opinions, to select specific AI-IP disputes of 38 claims based on relevant IP legal terms. Additionally, the study uses the AI and Copyright Case Tracker to include 21 world-leading AI copyright disputes, based on an AI-trained copyright legal analysis approach that refines the creation of a customized AI copyright case tracker for major international legal disputes, using specific initial analysis approaches involving AI copyright dispute analysis. The final dataset consisted of 131 cases: a base AI litigation dataset, a public court-opinion database, and an AI copyright case-tracking repository. These cases were selected after applying inclusion, exclusion, and deduplication procedures.
| Dataset Source | Dataset Type | No. of Cases Used | Coverage |
| Base AI Litigation Dataset1 | AI-related court cases | 72 (filtered subset) | U.S. federal and state courts |
| Caselaw Access Project (CAP)31 | Public legal case corpus | 38 (AI-IP focused) | U.S. court opinions |
| AI & Copyright Case Tracker (CMS)32 | Curated case summaries | 21 landmark cases | Global (US, EU, UK) |
Table 2: Dataset Description. Summary of datasets used for modeling intellectual property disputes in AI-generated content, including source, structure, and relevance to ownership analysis.
To ensure accuracy and consistency, a combination of manual and semi-automated processes was used. The labeling method used three annotators with backgrounds in both law and AI-related legal studies. All variables, including human creative input, AI system autonomy, training data legality, contractual ownership provisions, jurisdiction, transparency of AI involvement, and explicit decision rules and examples, were defined in an annotated codebook. Two annotators independently annotated each instance; differences were settled by discussion and, if needed, by a third expert reviewer. Cohen's kappa coefficient was used to measure inter-rater agreement; the result was 0.81, indicating strong agreement. About 9% of cases involved disagreements, mostly over AI autonomy and contract interpretation. For further Bayesian Network modeling, the completed annotations were utilized.
Identification of legal uncertainty factors
Key sources of uncertainty in AI-generated content ownership were identified through an analysis of AI-related intellectual property cases across multiple legal datasets, including litigation-based datasets and publicly available court repositories. Human creativity, AI autonomy, the legality of training data, contractual connections, jurisdictional issues, and the transparency of AI systems were important factors. The degree of human engagement and contractual ownership definitions were often highlighted by courts, and legal ambiguity was exacerbated by differences in copyright laws and the usage of copyrighted training data. The Bayesian Network framework for probabilistic ownership analysis was developed based on these considerations.
| Factor | Description |
| Human Creative Involvement | Degree of human contribution in generating the content (high, moderate, or minimal). |
| Level of AI Autonomy | Extent to which the AI system operates independently without human guidance. |
| Training Data Ownership and Legality | Legal status of the training data used (licensed, public domain, proprietary, or infringing). |
| Contractual Ownership Clauses | Existence and clarity of agreements defining ownership among stakeholders. |
| Jurisdictional Legal Framework | Applicable national or regional intellectual property laws. |
| Purpose of Use | Whether the AI-generated content is used for commercial or non-commercial purposes. |
| Transparency and Explainability | Availability of information regarding the AI system’s design and operation. |
| Disclosure of AI Assistance | Whether the use of AI in content creation was explicitly disclosed. |
| Nature of Generated Content | Type of output produced (e.g., text, image, code, or design). |
| Judicial Precedent Availability | Presence of prior court rulings relevant to AI-generated content ownership. |
Table 3: Key Legal and Technical Uncertainty Factors Influencing Ownership. Identification of critical variables affecting ownership attribution in AI-generated content, encompassing both legal and technical dimensions.
Table 3 summarizes the ownership attribution-related uncertainties in the proposed Bayesian Network model. These uncertainties are derived based on the intellectual property disputes associated with AI and the existing intellectual property laws. Human creative input and autonomy in AI refer to the relationship between human authors and AI-generated content in determining ownership. The uncertainties, including the legality of the training data used by AI, ownership of the contract, and the jurisdiction's existing intellectual property laws, are grounds for dispute. The purpose of use and AI assistance disclosure refers to the circumstances in which the judicial system views AI-assisted generation. Transparency and explainability, on the other hand, refer to the aspect associated with accountability and responsibility. Other uncertainties, including the nature of AI-generated creations and existing judicial precedents, determine how existing intellectual property laws apply to AI creations.
Data annotation and variable encoding
The dataset, constructed from filtered AI litigation cases, CAP, and selected AI copyright cases, among others, shall then be annotated manually and semi-automatically based on legal uncertainty factors identified for each case. For each case Structured labels are determined by evaluating court opinions, descriptions, and reasoning. The annotation task involves determining both legal and technical factors, such as the level of human creative participation, the autonomy of AI, the legal validity of the training data, contractual clauses on ownership, jurisdiction, and the disclosure of assistance by AI. For consistency, the guidelines are derived from other intellectual property principles and verified through cross-checking across a variety of legal sources. Refer to Supplementary File 1(S1) for a detailed mathematical description of data annotation and variable encoding.
The framework was implemented using a programming language with probabilistic graphical modeling libraries and numerical computing tools. Data preprocessing and analysis were performed in a notebook-based computing environment on a desktop system with a multi-core processor and standard memory configuration
To ensure consistency and reproducibility in dataset annotation, a structured variable-coding scheme was developed for all legal and technical factors included in the Bayesian Network framework. Each case was manually reviewed and annotated using predefined categorical states derived from legal precedents and AI governance literature. Table 4 summarizes the variables, descriptions, and coding categories used in the study. To facilitate probabilistic modeling and scenario-based inference within the Bayesian Network framework, the coding system was applied uniformly across all annotated cases. To represent often observed legal and technological circumstances in conflicts involving AI-generated material, variable states were chosen.
| Variable | Description | Coding Categories |
| Human Contribution | Degree of human creative involvement in content generation | Low, Medium, High |
| AI Autonomy | Level of independent AI system generation | Low, Medium, High |
| Training Data Legality | Legal status of datasets used for AI training | Licensed, Public Domain, Uncertain, Unauthorized |
| Contractual Ownership | Presence of contractual ownership clauses | Present, Absent |
| Attribution Clarity | Clarity of contributor identification | Clear, Partial, Unclear |
| Jurisdiction | Legal jurisdiction associated with the dispute | India, United States, China, Other |
| Transparency | Availability of information regarding AI generation process | High, Moderate, Low |
| Ownership Outcome | Final ownership attribution outcome (target variable) | Human-Owned, Shared Ownership, Organization-Owned, Uncertain |
Table 4: Variable Coding Scheme Used for Bayesian Network Annotation. The table summarizes the legal and technical variables used during dataset annotation, including their definitions and categorical coding states applied for probabilistic modeling and ownership inference.
Bayesian network structure design
The Bayesian Network (BN) was created to model the probabilistic relationships among contextual, technological, and legal factors that affect the ownership of information produced by artificial intelligence. While the target node depicted ownership results, the network nodes represented variables such as human contribution, AI autonomy, training data legality, contractual ownership terms, and jurisdiction. Conditional dependencies arising from data-driven associations and judicial reasoning processes were recorded by directed edges. Data-driven optimization and legal constraints were combined into a hybrid structure-learning method. Allowable dependencies were determined by judicial principles, and the most likely dependency graph explaining reported litigation results was found using score-based structure learning. This method ensured that the final network would remain both legally interpretable and statistically valid.

Figure 2: Bayesian Network structure used for ownership inference. Graphical representation of the Bayesian Network showing dependencies among legal and technical variables influencing ownership decisions. Please click here to view a larger version of this figure.
Figure 2 shows how a Bayesian Network is employed in AI-generated content ownership assignments. On a top-level model, some essential input variables, namely Human Input, Training Data, AI Autonomy, and Jurisdiction, indicate some essential legal and technical dimensions that should be considered when rendering a right over AI-generated content. Every variable is designated as a probable node rather than a hard-coded rule, enabling a model that accounts for variable dependencies. The green nodes, also known as connection points, appear between the input nodes and the final output. They indicate, for instance, how levels of human involvement correlate with AI autonomy, while both legal training data and jurisdictional factors help define a right to AI-generated content from a jurisdictional perspective.
Legal theory and annotated litigation data served as the foundation for the Bayesian Network's variable selection and discrete state definitions. Given their importance in determining ownership, key factors such as human involvement, AI autonomy, the legality of training data, contractual conditions, jurisdiction, and transparency were given top priority. Expert-reviewed annotation criteria were used to discretize variables into legally significant categories. To identify statistically supported dependencies, a hybrid strategy combining score-based learning and legal constraints was used to construct the network structure. The approach was further validated using sensitivity and ablation studies, which confirmed the significant impact of human input and contractual restrictions on ownership inference, while other variables showed moderating effects. For decision-support systems, this combined legal-empirical approach enhances both robustness and interpretability.
The inferred probabilities from all interconnected nodes converge at the Ownership Decision node, which generates posterior ownership probabilities based on the available legal and technical evidence. This enables ownership inference even under incomplete or disputed conditions and supports dynamic updating when new evidence becomes available. To reduce overfitting and spurious correlations, model selection and sensitivity analysis were used to remove redundant or low-impact edges, yielding a concise directed acyclic graph (DAG). The final Bayesian Network structure captures both probabilistic dependencies and causal relationships among factors contributing to legal uncertainty in AI-generated content ownership. The mathematical description of the Bayesian network design is provided in Supplementary File 1 (S2).
The Bayesian Network enables probabilistic ownership inference even when some legal or technical factors are incomplete or partially observable. By modifying evidence variables, the framework also supports scenario-based analysis for legal decision-making, policy evaluation, and AI governance applications. The network structure was developed using a hybrid approach combining expert-defined legal constraints with data-driven learning. A Hill Climbing algorithm with Bayesian Information Criterion (BIC) scoring was used for structure optimization, while Bayesian parameter estimation with Dirichlet priors and Laplace smoothing (&α = 1) addressed data sparsity and missing observations. All variables were represented as discrete categorical states to support probabilistic inference. Python probabilistic graphical modeling modules were used to construct the Bayesian Network framework. The two phases of network building were parameter learning and structure definition. Initially, the network's nodes represented the legal and technical characteristics found during dataset annotation. A hybrid method that coupled data-driven structural dependency analysis with expert-defined legal relationships was used to create directed edges between nodes. Maximum likelihood estimation was used to estimate Conditional Probability Tables (CPTs) from the annotated dataset. The probabilistic graphical modeling techniques within a computational environment package were used to implement the Bayesian Network, with the Maximum Likelihood Estimator function for parameter learning and the Bayesian Network () function for defining the network structure. The Variable Elimination () inference procedure was used to calculate posterior ownership probabilities under various evidence conditions.
Parameter learning
In quantitatively modeling legal uncertainty in AI-generated content ownership, parameter estimation is a crucial step in the proposed Bayesian Network framework. Bayesian parameter estimate was used to learn Conditional Probability Tables (CPTs) from annotated legal cases. While the goal node indicates ownership attribution, each node reflects a legal or technological aspect, such as human contribution, AI autonomy, the legality of training data, or contractual ownership provisions. Because Bayesian estimates can handle scant legal material and inadequate evidence, it was chosen to promote interpretable legal reasoning and to represent probabilistic correlations observed in judicial decisions. Representative conditional probability values acquired within the Bayesian Network are shown in Table 5. While great AI autonomy promotes uncertainty or no-ownership outcomes, significant human creative engagement indicates a substantial probability of human ownership, indicating judicial emphasis on human intellectual contribution. The findings also emphasize the importance of contractual ownership terms and the legality of training data, with explicit agreements and licensed data increasing ownership. The framework's sensitivity to various legal situations is further demonstrated by jurisdictional variances.
| Parent Variables | States of Parent Variables | Ownership Outcome | Learned Conditional Probability |
| Human Creative Involvement | High | Human Author | 0.82 |
| Human Creative Involvement | Low | AI Developer | 0.41 |
| AI Autonomy | High | No Ownership Recognized | 0.55 |
| Training Data Legality | Licensed | Human / Developer | 0.76 |
| Training Data Legality | Infringing | No Ownership / Disputed | 0.68 |
| Contractual Ownership Clause | Explicit Present | Contractual Party | 0.89 |
| Contractual Ownership Clause | Absent | Disputed / Uncertain | 0.61 |
| Jurisdictional Framework | Human-Centric IP Law | Human Author | 0.79 |
| Jurisdictional Framework | Ambiguous / Emerging Law | Joint / Uncertain | 0.47 |
Table 5: Example Conditional Probabilities in the Bayesian Network Model. Illustrative conditional probability values learned by the Bayesian Network, demonstrating dependencies among key variables influencing ownership outcomes.
Algorithm 1 encapsulates the entire process of the proposed Bayesian Network framework for capturing ownership uncertainties in AI content. Algorithm 1: Initiate a Bayesian Network structure by defining the graphical model where the nodes are the technical and legal uncertainties examined in the litigations with annotated information, and the target node is the ownership designation. The structure of the Bayesian Network model can then be established by applying a joint approach that merges the existing legally constraining structures with optimized structures derived from the data. The conditional probability tables can then be specified by implementing the Bayesian parameter estimation scheme based on the obtained dataset, thereby accounting for uncertainties, sparsity, and conditional relationships inferred by the judiciary. When applying the algorithm to a particular legal or hypothetical example, the input information or evidence can be elicited in the Bayesian Network by applying probabilistic inference to determine the posterior probability of ownership. The algorithm can then apply adjustments to the input values by iteratively modifying the evidence input to the Bayesian Network and assessing the impact on ownership designation by adjusting the human input, AI autonomy, contractual legality, and the legality of the training data. Refer to Supplementary File 1(S3) for Algorithmic steps.
Probabilistic inference and scenario analysis
The Bayesian Network's graphical structure and conditional probability tables can then be trained on an annotated dataset of AI-related litigation regarding ownership, and the system enables inference on the likelihood of ownership. The mathematical descriptions are stated in Supplementary File 1 (S4).
By observing how these posterior ownership probabilities vary across counterfactual scenarios, the model yields a set of relative influences on judicial outcomes for each legal and technical feature. This scenario-driven inference capability constitutes one of the main contributions of the proposed work, since it aligns with how courts, policymakers, and legal experts reason about the ownership of AI-generated content: through contextual evaluation rather than the rigid application of rules. The Bayesian Network hereby serves as an explainable decision-support tool, transforming empirical litigation patterns into probabilistic insights that support consistent reasoning across jurisdictions, contractual contexts, and levels of AI autonomy. Accordingly, such a usability factorial moves beyond descriptive legal analysis by providing a systematic mechanism for anticipating ownership outcomes under evolving AI and intellectual property regimes. The Variable Elimination approach was used to calculate the posterior probability of ownership outcomes for inference. Iteratively changing the evidence variables and recalculating posterior distributions allowed for scenario analysis. To ensure resilience across small datasets, the model was validated using leave-one-out validation and k-fold cross-validation (k = 5). To perform a sensitivity analysis, input variables were systematically varied, and the accompanying changes in posterior probabilities (ΔP(Y)) were measured.
Model validation and sensitivity analysis
The suggested Bayesian Network model was verified using an annotated dataset of AI-related intellectual property conflicts to guarantee dependability, robustness, and legal applicability. To forecast ownership outcomes and compare them with actual legal decisions, a retrospective validation approach was used. Given the scarcity of AI-related ownership dispute situations, leave-one-out and k-fold cross-validation were used to assess the model's robustness.
The impact of legal and technical factors on ownership attribution, including human involvement, AI autonomy, the legality of training data, and contractual ownership terms, was investigated through a sensitivity analysis. The investigation identified contextual and high-impact factors that influence ownership choices by altering parent node states and monitoring changes in posterior probabilities. By showing how legal ambiguity spreads throughout the network, this enhanced interpretability. The validation and sensitivity analysis findings are collected in Table 6. During cross-validation, the model achieved more than 80% agreement between inferred and actual judicial outcomes. Robustness testing verified trustworthy inference under limited-evidence settings, and low posterior variance across folds suggested stable, generalizable conditional probabilities. Sensitivity analysis revealed that while AI autonomy and the legality of training data greatly increased attribution uncertainty, contractual ownership clauses and human creative involvement had the greatest impact on ownership attribution. Technical and contractual conditions interacted with jurisdictional variables, exerting a moderate effect. Overall, the findings show that the suggested paradigm is both legally interpretable and empirically sound.
| Validation / Sensitivity Aspect | Method Applied | Metric / Observation |
| Case-Based Validation | Leave-one-out cross-validation | 81.3% ownership outcome consistency |
| Cross-Fold Stability | 5-fold cross-validation | Posterior variance < 0.05 across folds |
| Missing Evidence Robustness | Partial evidence inference | < 7% change in posterior probabilities |
| Sensitivity: Human Creative Involvement | One-factor perturbation | ΔP(Y) ≈ +0.32 |
| Sensitivity: AI Autonomy | One-factor perturbation | ΔP(Y) ≈ −0.28 |
| Sensitivity: Training Data Legality | One-factor perturbation | ΔP(Y) ≈ +0.25 |
| Sensitivity: Contractual Clauses | One-factor perturbation | ΔP(Y) ≈ +0.37 |
| Sensitivity: Jurisdiction | One-factor perturbation | ΔP(Y) ≈ +0.14 |
Table 6: Validation and Sensitivity Analysis Results. Results of model validation and sensitivity analysis, evaluating robustness, stability, and the impact of key variables on ownership inference.
Interpretation and policy implication analysis
The probabilistic outputs of the Bayesian Network provide an interpretable framework for analyzing ownership of AI-generated content by modeling legal and technical factors, including human contribution, AI autonomy, contractual clarity, and the legality of the training data. Rather than serving solely as predictive outputs, the posterior probability estimates support transparent, scenario-based reasoning under uncertain legal conditions. The framework can assist legal professionals in improving consistency during ownership evaluation, support policymakers in identifying regulatory gaps related to AI-generated content, and help AI developers assess how design and governance choices may influence ownership attribution in different legal contexts.