Research Article

Modeling Legal Uncertainty in AI-Generated Content Ownership: A Bayesian Network Approach to Intellectual Property Rights Allocation

DOI:

10.3791/71101

June 12th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study presents a Bayesian Network–based workflow for analyzing ownership of AI-generated content under legal and technical uncertainty. Using AI litigation, copyright datasets, and framework models, factors such as human contribution, contracts, legality of training data, and AI autonomy to support transparent, interpretable, and scenario-based decision-making in intellectual property disputes.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The rapid adoption of generative artificial intelligence has exposed limitations in traditional intellectual property frameworks based on human authorship, creating uncertainty in ownership attribution. This study proposes a Bayesian Network–based workflow for analyzing ownership of AI-generated content under legal and technical uncertainty. The workflow integrates dataset curation, variable annotation, probabilistic dependency modeling, inference generation, cross-validation, and sensitivity analysis to evaluate relationships among factors, including human contribution, the legality of training data, contractual context, and AI autonomy. The framework was developed using annotated cases derived from a curated AI litigation dataset, publicly available court opinion databases, and AI copyright case trackers. A sensitivity analysis was conducted to examine the influence of legal and technical variables on ownership attribution across different scenarios. The proposed framework supports probabilistic and interpretable reasoning for AI-related intellectual property disputes and provides a structured decision-support approach for legal professionals, policymakers, and AI developers. Results demonstrated stable posterior probability estimates across folds, with ownership prediction consistency exceeding 80% across datasets. The framework enables probabilistic, scenario-based reasoning and provides an interpretable decision-support tool, improving transparency and consistency in resolving ownership disputes for legal professionals, policymakers, and AI developers.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The growing use of artificial intelligence (AI) systems in high-stakes industries, including healthcare, criminal justice, finance, and digital platforms, raises important questions about legal responsibility, accountability, and fairness. Despite rapid progress, a large portion of the literature on AI governance is either normative or predictive and lacks empirical support1,2. Current methods for discovering AI-related problems, such as event databases, survey-based research, and ethical frameworks, frequently fall short of capturing persistent problems in the actual world. These approaches either overemphasize high-profile occurrences while ignoring systemic, practice-level issues or reflect expectations rather than observable damage3,4. The majority of incident databases rely on voluntary reporting and media coverage, which frequently highlights high-profile failures while underrepresenting common but important AI-related problems5,6. By enabling systems to produce text, images, and music using transformer-based deep learning architectures, generative AI, especially Large Language Models (LLMs), has made content production even more challenging7,8,9. Traditional intellectual property frameworks based on human authorship are facing difficulties as a result of this change, making it harder to distinguish between human and machine-generated innovation10,11,12.

Previous studies on AI-generated content and intellectual property have mostly focused on policy analysis and legal interpretation, including conceptual discussions of authorship, originality, and ownership13. Even while these studies offer valuable insights, they remain primarily descriptive and lack computational methods to address ownership-determination uncertainty14. Simultaneously, technological methods to support content attribution and provenance have been developed, including watermarking, fingerprinting, and dataset protection15. The legal distribution of ownership, however, is not addressed by these approaches, especially in multi-stakeholder settings with developers, users, and data producers. Additionally, current interdisciplinary studies have not made sufficient progress toward probabilistic models that can integrate technical and legal factors for ownership evaluation in the face of uncertainty16,17,18,19,20. Legal studies have highlighted the limitations of existing intellectual property frameworks in addressing ownership of AI-generated content. In many jurisdictions, including under Section 13 of the Indian Copyright Act, AI-generated works often fail to satisfy traditional authorship criteria, creating ownership ambiguity21,22. Although prior studies discuss judicial interpretations and copyright implications, they lack formal mechanisms for modeling uncertainty in ownership determination22,23. To address this gap, this study proposes a Bayesian Network–based framework that models probabilistic relationships among human involvement, training data provenance, contractual context, and AI autonomy for scenario-based ownership analysis24.

Research has explored the application of machine learning models such as logistic regression, support vector machines (SVMs), convolutional neural networks (CNNs), and long short-term memory (LSTM) networks in legal decision-making, demonstrating their usefulness in high-volume legal environments where manual precedent analysis is time-consuming25. However, these approaches often lack interpretability and transparency, which are essential for accountable judicial decision support. Consequently, Explainable Artificial Intelligence (XAI) methods have attracted attention for improving interpretability and providing understandable reasoning for AI-assisted decisions26,27,28.

Methodology / ApproachData UsedProsCons / Gaps
Review of legal & technical instruments to link content to training data & propose frameworks for attribution and accountability.29Literature survey, case examplesIntegrates tech + law perspective; proposes actionable attribution strategiesLacks formal probabilistic modeling; no quantitative validation
Taxonomy of technical IP protection methods for training/outputs of generative models.30Survey of technical approaches and risksSystematic review; organizes existing techniques for IP protectionFocuses on training-data protection vs ownership of generated output
Proposes latent fingerprinting / watermarking to verify AI origin & IP attribution.31Experiments with generative models (text & image)Practical verification tool; strong for ownership evidenceWatermarking doesn’t resolve legal ownership frameworks itself
Comparative legal analysis of authorship & ownership standards across jurisdictions.32Legal doctrine, statutory texts, representative casesExplains how human contribution affects ownership decisionsDescriptive; does not propose computational models

Table 1: Summary of Methodologies, Data Sources, and Key Findings in Prior Work. Overview of existing studies on AI-generated content ownership, highlighting methodological approaches, datasets utilized, and major findings.

Table 1 summarizes the existing methodologies, data sources, and key findings29,30,31,32. This table discusses the ownership of AI-generated content in the literature and highlights the methodological approaches, utilized datasets, and the findings of the considered literature. Recent studies on AI governance and copyright disputes have further highlighted challenges related to authorship, training data legality, and ownership allocation in AI-generated content33,34. Although approaches such as provenance tracking and evidence validation contribute to attribution analysis34, existing rule-based and machine learning methods remain limited in handling ambiguity, conflicting evidence, and legally interpretable ownership reasoning. Technical approaches such as watermarking and provenance analysis also primarily focus on attribution rather than ownership determination.

This paper proposes a Bayesian Network-based framework for evaluating ownership of AI-generated content amid legal and technical ambiguity, overcoming the shortcomings of current deterministic, rule-based approaches. The approach enables scenario-based inference from incomplete or ambiguous evidence by modeling probabilistic interactions among key aspects, including human engagement, training data provenance, contractual context, and AI autonomy. The suggested method, unlike traditional predictive models, offers interpretable posterior probability estimates that facilitate clear, organized reasoning consistent with judicial decision-making procedures.

Important legal and technological concepts are explicitly defined to ensure consistency throughout the study. While “ownership” relates to the distribution of rights related to the produced work, “authorship” refers to the legal acknowledgment of a creator. “Attribution” may not always indicate legal ownership, but it does identify contributing entities. “Explainability” refers to a model's capacity to provide a comprehensible justification for probabilistic outcomes, whereas “legal uncertainty” refers to ambiguity arising from incomplete or changing legal standards. By incorporating legal and technical factors into an interpretable probabilistic workflow, the suggested approach aims to facilitate structured decision-making in AI-related intellectual property conflicts. For legal experts, legislators, and AI developers assessing ownership concerns in generative AI systems, the technique offers a computational viewpoint.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study does not involve human participants or animal subjects. The analysis is based solely on publicly available legal case materials and curated legal datasets. Therefore, ethics approval and informed consent were not required.

The proposed methodology aims to capture the legal uncertainty surrounding ownership of AI-created content by developing a Bayesian Network based on actual case data. To start with, cases related to AI, including copyright, patent, and trade secrets, would be harvested from an existing corpus and each tagged with actual legal outcome data. Based upon legal precedent and existing case studies, uncertainty parameters that could have an impact on content ownership would include creative human input, AI system autonomy, legality of training data, contractual definitions of ownership, jurisdiction, and transparency, each assigned as nodes within the Bayesian Network with discreet states indicating differing legal and technical realities, and with content ownership determination as the target node. The structure of this network is discovered by integrating legal expertise with data-driven structural discovery to promote both interpretability and legal plausibility. The tables of conditional probabilities are then approximated based on litigation outcomes identified in cases, allowing the network to process uncertain and incomplete data probabilistically. Subsequent to model training, it allows for both inferencing and scenario testing to determine how changes in contributing factors impact inferred ownership in AI-generated content based on this constructed Bayesian Network, followed by testing and sensitivity analysis validation prior to inferring results that facilitate future legal implications for intellectual property in AI applications by developers and in courts. Figure 1 shows the architectural workflow of the proposed work.

Bayesian network model diagram for legal case analysis using AI; includes data pre-processing and simulations.
Figure 1: Architectural Diagram of the Proposed Framework: Overview of the proposed Bayesian Network–based system for determining ownership of AI-generated content. The architecture illustrates the flow from data acquisition and annotation to probabilistic modeling, inference, and ownership attribution. Please click here to view a larger version of this figure.

Figure 1 illustrates the architectural workflow of the proposed Bayesian Network–based framework for AI-generated content ownership analysis. The workflow includes data collection, feature annotation, Bayesian Network modeling, probabilistic inference, scenario analysis, and validation. Legal and technical variables, such as human involvement, AI autonomy, training data provenance, contractual context, and jurisdiction, are modeled within a Bayesian framework to estimate ownership attribution under uncertain conditions. The framework supports interpretable probabilistic reasoning and scenario-based evaluation for legal decision support.

Data collection
Data on legal cases was collected from various sources to create a dataset focused on ownership and intellectual property lawsuits involving AI-generated content. To start, the general dataset of legal cases sourced in the base paper was filtered to include only litigations concerning copyright, ownership, IP-related disputes, and other IP matters in the training data. Additional legal cases were obtained from publicly available court-opinion databases that provide access to court opinions from millions of cases, as well as a copyright case-tracking repository that tracks both ongoing and concluded legal proceedings worldwide on uses of Artificial Intelligence regarding Intellectual Property. Legal cases were also annotated to highlight key variables, including the legal case outcome, the level of human creative input, the autonomy of the Artificial Intelligence system, the origin of the training data, jurisdiction, and contract parameters, if applicable. This is a multi-source dataset that aggregates sufficient data on Intellectual Property related to the use of AI tools, suitable for probabilistic reasoning in Bayesian network analysis.

Three main sources, including the base AI litigation dataset, a public court-opinion database, and an AI copyright case tracking repository, were used in a methodical collection and screening process to create the dataset for this study. Keywords such as “artificial intelligence,” “AI-generated content,” “copyright,” “intellectual property,” “authorship, “ownership,” and “machine-generated works” were used throughout repository search interfaces and metadata fields to find pertinent cases. Cases involving AI-generated or AI-assisted content that addressed ownership, authorship, or intellectual property concerns with enough depth for variable extraction were considered. Cases that were duplicates, unconnected to AI-generated content, concentrated on other legal areas, or lacked adequate information were not accepted. A two-phase screening procedure was used, with a full-text eligibility assessment coming after the initial title and summary examination. To ensure uniqueness, duplicate records across multiple datasets were identified using case metadata and removed.

The keywords “artificial intelligence,” “AI-generated content,” “machine-generated works,” “copyright,” “authorship,”“ownership,”“training data,” and “intellectual property dispute” were used in the dataset search process. To guarantee the quality and relevance of the dataset, a two-stage screening procedure was used. To find ownership and copyright conflicts pertaining to AI, titles and case summaries were reviewed in the first stage. To determine eligibility for inclusion in the study, a full-text case review was conducted in the second stage. Only cases involving AI-generated or AI-assisted content, with adequate legal and technical details regarding authorship, ownership attribution, contractual context, or training data usage, were included. Excluded were cases with insufficient contextual information, inadequate legal documents, duplicate entries, and cases unrelated to intellectual property conflicts. By comparing metadata such as case title, jurisdiction, filing date, and dispute category, duplicate cases across sources were identified and eliminated. Annotation, Bayesian Network modeling, probabilistic inference, and sensitivity analysis were then performed on the final curated dataset.

The data employed in the study comprises a compilation of legal sources covering the full scope of disputes over AI content ownership rights and intellectual property claims, as shown in Table 2. First, the AI litigation data used in the primary study includes the selective preservation of about 72 legal disputes specifically tied to copyright claims, authorship, authorship by invention, and ownership involving AI outputs. To increase data variability for better representation of emerging disputes in the realm of generative AI, the data includes the acquisition of relevant disputes sourced from the publicly available court opinion database (CAP), a massive publicly available collection of U.S. court opinions, to select specific AI-IP disputes of 38 claims based on relevant IP legal terms. Additionally, the study uses the AI and Copyright Case Tracker to include 21 world-leading AI copyright disputes, based on an AI-trained copyright legal analysis approach that refines the creation of a customized AI copyright case tracker for major international legal disputes, using specific initial analysis approaches involving AI copyright dispute analysis. The final dataset consisted of 131 cases: a base AI litigation dataset, a public court-opinion database, and an AI copyright case-tracking repository. These cases were selected after applying inclusion, exclusion, and deduplication procedures.

Dataset SourceDataset TypeNo. of Cases UsedCoverage
Base AI Litigation Dataset1AI-related court cases72 (filtered subset)U.S. federal and state courts
Caselaw Access Project (CAP)31Public legal case corpus38 (AI-IP focused)U.S. court opinions
AI & Copyright Case Tracker (CMS)32Curated case summaries21 landmark casesGlobal (US, EU, UK)

Table 2: Dataset Description. Summary of datasets used for modeling intellectual property disputes in AI-generated content, including source, structure, and relevance to ownership analysis.

To ensure accuracy and consistency, a combination of manual and semi-automated processes was used. The labeling method used three annotators with backgrounds in both law and AI-related legal studies. All variables, including human creative input, AI system autonomy, training data legality, contractual ownership provisions, jurisdiction, transparency of AI involvement, and explicit decision rules and examples, were defined in an annotated codebook. Two annotators independently annotated each instance; differences were settled by discussion and, if needed, by a third expert reviewer. Cohen's kappa coefficient was used to measure inter-rater agreement; the result was 0.81, indicating strong agreement. About 9% of cases involved disagreements, mostly over AI autonomy and contract interpretation. For further Bayesian Network modeling, the completed annotations were utilized.

Identification of legal uncertainty factors
Key sources of uncertainty in AI-generated content ownership were identified through an analysis of AI-related intellectual property cases across multiple legal datasets, including litigation-based datasets and publicly available court repositories. Human creativity, AI autonomy, the legality of training data, contractual connections, jurisdictional issues, and the transparency of AI systems were important factors. The degree of human engagement and contractual ownership definitions were often highlighted by courts, and legal ambiguity was exacerbated by differences in copyright laws and the usage of copyrighted training data. The Bayesian Network framework for probabilistic ownership analysis was developed based on these considerations.

FactorDescription
Human Creative InvolvementDegree of human contribution in generating the content (high, moderate, or minimal).
Level of AI AutonomyExtent to which the AI system operates independently without human guidance.
Training Data Ownership and LegalityLegal status of the training data used (licensed, public domain, proprietary, or infringing).
Contractual Ownership ClausesExistence and clarity of agreements defining ownership among stakeholders.
Jurisdictional Legal FrameworkApplicable national or regional intellectual property laws.
Purpose of UseWhether the AI-generated content is used for commercial or non-commercial purposes.
Transparency and ExplainabilityAvailability of information regarding the AI system’s design and operation.
Disclosure of AI AssistanceWhether the use of AI in content creation was explicitly disclosed.
Nature of Generated ContentType of output produced (e.g., text, image, code, or design).
Judicial Precedent AvailabilityPresence of prior court rulings relevant to AI-generated content ownership.

Table 3: Key Legal and Technical Uncertainty Factors Influencing Ownership. Identification of critical variables affecting ownership attribution in AI-generated content, encompassing both legal and technical dimensions.

Table 3 summarizes the ownership attribution-related uncertainties in the proposed Bayesian Network model. These uncertainties are derived based on the intellectual property disputes associated with AI and the existing intellectual property laws. Human creative input and autonomy in AI refer to the relationship between human authors and AI-generated content in determining ownership. The uncertainties, including the legality of the training data used by AI, ownership of the contract, and the jurisdiction's existing intellectual property laws, are grounds for dispute. The purpose of use and AI assistance disclosure refers to the circumstances in which the judicial system views AI-assisted generation. Transparency and explainability, on the other hand, refer to the aspect associated with accountability and responsibility. Other uncertainties, including the nature of AI-generated creations and existing judicial precedents, determine how existing intellectual property laws apply to AI creations.

Data annotation and variable encoding
The dataset, constructed from filtered AI litigation cases, CAP, and selected AI copyright cases, among others, shall then be annotated manually and semi-automatically based on legal uncertainty factors identified for each case. For each case Structured labels are determined by evaluating court opinions, descriptions, and reasoning. The annotation task involves determining both legal and technical factors, such as the level of human creative participation, the autonomy of AI, the legal validity of the training data, contractual clauses on ownership, jurisdiction, and the disclosure of assistance by AI. For consistency, the guidelines are derived from other intellectual property principles and verified through cross-checking across a variety of legal sources. Refer to Supplementary File 1(S1) for a detailed mathematical description of data annotation and variable encoding.

The framework was implemented using a programming language with probabilistic graphical modeling libraries and numerical computing tools. Data preprocessing and analysis were performed in a notebook-based computing environment on a desktop system with a multi-core processor and standard memory configuration

To ensure consistency and reproducibility in dataset annotation, a structured variable-coding scheme was developed for all legal and technical factors included in the Bayesian Network framework. Each case was manually reviewed and annotated using predefined categorical states derived from legal precedents and AI governance literature. Table 4 summarizes the variables, descriptions, and coding categories used in the study. To facilitate probabilistic modeling and scenario-based inference within the Bayesian Network framework, the coding system was applied uniformly across all annotated cases. To represent often observed legal and technological circumstances in conflicts involving AI-generated material, variable states were chosen.

VariableDescriptionCoding Categories
Human ContributionDegree of human creative involvement in content generationLow, Medium, High
AI AutonomyLevel of independent AI system generationLow, Medium, High
Training Data LegalityLegal status of datasets used for AI trainingLicensed, Public Domain, Uncertain, Unauthorized
Contractual OwnershipPresence of contractual ownership clausesPresent, Absent
Attribution ClarityClarity of contributor identificationClear, Partial, Unclear
JurisdictionLegal jurisdiction associated with the disputeIndia, United States, China, Other
TransparencyAvailability of information regarding AI generation processHigh, Moderate, Low
Ownership OutcomeFinal ownership attribution outcome (target variable)Human-Owned, Shared Ownership, Organization-Owned, Uncertain

Table 4: Variable Coding Scheme Used for Bayesian Network Annotation. The table summarizes the legal and technical variables used during dataset annotation, including their definitions and categorical coding states applied for probabilistic modeling and ownership inference.

Bayesian network structure design
The Bayesian Network (BN) was created to model the probabilistic relationships among contextual, technological, and legal factors that affect the ownership of information produced by artificial intelligence. While the target node depicted ownership results, the network nodes represented variables such as human contribution, AI autonomy, training data legality, contractual ownership terms, and jurisdiction. Conditional dependencies arising from data-driven associations and judicial reasoning processes were recorded by directed edges. Data-driven optimization and legal constraints were combined into a hybrid structure-learning method. Allowable dependencies were determined by judicial principles, and the most likely dependency graph explaining reported litigation results was found using score-based structure learning. This method ensured that the final network would remain both legally interpretable and statistically valid.

AI governance diagram shows factors influencing ownership decisions: human input, autonomy, data.
Figure 2: Bayesian Network structure used for ownership inference. Graphical representation of the Bayesian Network showing dependencies among legal and technical variables influencing ownership decisions. Please click here to view a larger version of this figure.

Figure 2 shows how a Bayesian Network is employed in AI-generated content ownership assignments. On a top-level model, some essential input variables, namely Human Input, Training Data, AI Autonomy, and Jurisdiction, indicate some essential legal and technical dimensions that should be considered when rendering a right over AI-generated content. Every variable is designated as a probable node rather than a hard-coded rule, enabling a model that accounts for variable dependencies. The green nodes, also known as connection points, appear between the input nodes and the final output. They indicate, for instance, how levels of human involvement correlate with AI autonomy, while both legal training data and jurisdictional factors help define a right to AI-generated content from a jurisdictional perspective.

Legal theory and annotated litigation data served as the foundation for the Bayesian Network's variable selection and discrete state definitions. Given their importance in determining ownership, key factors such as human involvement, AI autonomy, the legality of training data, contractual conditions, jurisdiction, and transparency were given top priority. Expert-reviewed annotation criteria were used to discretize variables into legally significant categories. To identify statistically supported dependencies, a hybrid strategy combining score-based learning and legal constraints was used to construct the network structure. The approach was further validated using sensitivity and ablation studies, which confirmed the significant impact of human input and contractual restrictions on ownership inference, while other variables showed moderating effects. For decision-support systems, this combined legal-empirical approach enhances both robustness and interpretability.

The inferred probabilities from all interconnected nodes converge at the Ownership Decision node, which generates posterior ownership probabilities based on the available legal and technical evidence. This enables ownership inference even under incomplete or disputed conditions and supports dynamic updating when new evidence becomes available. To reduce overfitting and spurious correlations, model selection and sensitivity analysis were used to remove redundant or low-impact edges, yielding a concise directed acyclic graph (DAG). The final Bayesian Network structure captures both probabilistic dependencies and causal relationships among factors contributing to legal uncertainty in AI-generated content ownership. The mathematical description of the Bayesian network design is provided in Supplementary File 1 (S2).

The Bayesian Network enables probabilistic ownership inference even when some legal or technical factors are incomplete or partially observable. By modifying evidence variables, the framework also supports scenario-based analysis for legal decision-making, policy evaluation, and AI governance applications. The network structure was developed using a hybrid approach combining expert-defined legal constraints with data-driven learning. A Hill Climbing algorithm with Bayesian Information Criterion (BIC) scoring was used for structure optimization, while Bayesian parameter estimation with Dirichlet priors and Laplace smoothing (&α = 1) addressed data sparsity and missing observations. All variables were represented as discrete categorical states to support probabilistic inference. Python probabilistic graphical modeling modules were used to construct the Bayesian Network framework. The two phases of network building were parameter learning and structure definition. Initially, the network's nodes represented the legal and technical characteristics found during dataset annotation. A hybrid method that coupled data-driven structural dependency analysis with expert-defined legal relationships was used to create directed edges between nodes. Maximum likelihood estimation was used to estimate Conditional Probability Tables (CPTs) from the annotated dataset. The probabilistic graphical modeling techniques within a computational environment package were used to implement the Bayesian Network, with the Maximum Likelihood Estimator function for parameter learning and the Bayesian Network () function for defining the network structure. The Variable Elimination () inference procedure was used to calculate posterior ownership probabilities under various evidence conditions.

Parameter learning
In quantitatively modeling legal uncertainty in AI-generated content ownership, parameter estimation is a crucial step in the proposed Bayesian Network framework. Bayesian parameter estimate was used to learn Conditional Probability Tables (CPTs) from annotated legal cases. While the goal node indicates ownership attribution, each node reflects a legal or technological aspect, such as human contribution, AI autonomy, the legality of training data, or contractual ownership provisions. Because Bayesian estimates can handle scant legal material and inadequate evidence, it was chosen to promote interpretable legal reasoning and to represent probabilistic correlations observed in judicial decisions. Representative conditional probability values acquired within the Bayesian Network are shown in Table 5. While great AI autonomy promotes uncertainty or no-ownership outcomes, significant human creative engagement indicates a substantial probability of human ownership, indicating judicial emphasis on human intellectual contribution. The findings also emphasize the importance of contractual ownership terms and the legality of training data, with explicit agreements and licensed data increasing ownership. The framework's sensitivity to various legal situations is further demonstrated by jurisdictional variances.

Parent VariablesStates of Parent VariablesOwnership OutcomeLearned Conditional Probability
Human Creative InvolvementHighHuman Author0.82
Human Creative InvolvementLowAI Developer0.41
AI AutonomyHighNo Ownership Recognized0.55
Training Data LegalityLicensedHuman / Developer0.76
Training Data LegalityInfringingNo Ownership / Disputed0.68
Contractual Ownership ClauseExplicit PresentContractual Party0.89
Contractual Ownership ClauseAbsentDisputed / Uncertain0.61
Jurisdictional FrameworkHuman-Centric IP LawHuman Author0.79
Jurisdictional FrameworkAmbiguous / Emerging LawJoint / Uncertain0.47

Table 5: Example Conditional Probabilities in the Bayesian Network Model. Illustrative conditional probability values learned by the Bayesian Network, demonstrating dependencies among key variables influencing ownership outcomes.

Algorithm 1 encapsulates the entire process of the proposed Bayesian Network framework for capturing ownership uncertainties in AI content. Algorithm 1: Initiate a Bayesian Network structure by defining the graphical model where the nodes are the technical and legal uncertainties examined in the litigations with annotated information, and the target node is the ownership designation. The structure of the Bayesian Network model can then be established by applying a joint approach that merges the existing legally constraining structures with optimized structures derived from the data. The conditional probability tables can then be specified by implementing the Bayesian parameter estimation scheme based on the obtained dataset, thereby accounting for uncertainties, sparsity, and conditional relationships inferred by the judiciary. When applying the algorithm to a particular legal or hypothetical example, the input information or evidence can be elicited in the Bayesian Network by applying probabilistic inference to determine the posterior probability of ownership. The algorithm can then apply adjustments to the input values by iteratively modifying the evidence input to the Bayesian Network and assessing the impact on ownership designation by adjusting the human input, AI autonomy, contractual legality, and the legality of the training data. Refer to Supplementary File 1(S3) for Algorithmic steps.

Probabilistic inference and scenario analysis
The Bayesian Network's graphical structure and conditional probability tables can then be trained on an annotated dataset of AI-related litigation regarding ownership, and the system enables inference on the likelihood of ownership. The mathematical descriptions are stated in Supplementary File 1 (S4).

By observing how these posterior ownership probabilities vary across counterfactual scenarios, the model yields a set of relative influences on judicial outcomes for each legal and technical feature. This scenario-driven inference capability constitutes one of the main contributions of the proposed work, since it aligns with how courts, policymakers, and legal experts reason about the ownership of AI-generated content: through contextual evaluation rather than the rigid application of rules. The Bayesian Network hereby serves as an explainable decision-support tool, transforming empirical litigation patterns into probabilistic insights that support consistent reasoning across jurisdictions, contractual contexts, and levels of AI autonomy. Accordingly, such a usability factorial moves beyond descriptive legal analysis by providing a systematic mechanism for anticipating ownership outcomes under evolving AI and intellectual property regimes. The Variable Elimination approach was used to calculate the posterior probability of ownership outcomes for inference. Iteratively changing the evidence variables and recalculating posterior distributions allowed for scenario analysis. To ensure resilience across small datasets, the model was validated using leave-one-out validation and k-fold cross-validation (k = 5). To perform a sensitivity analysis, input variables were systematically varied, and the accompanying changes in posterior probabilities (ΔP(Y)) were measured.

Model validation and sensitivity analysis
The suggested Bayesian Network model was verified using an annotated dataset of AI-related intellectual property conflicts to guarantee dependability, robustness, and legal applicability. To forecast ownership outcomes and compare them with actual legal decisions, a retrospective validation approach was used. Given the scarcity of AI-related ownership dispute situations, leave-one-out and k-fold cross-validation were used to assess the model's robustness.

The impact of legal and technical factors on ownership attribution, including human involvement, AI autonomy, the legality of training data, and contractual ownership terms, was investigated through a sensitivity analysis. The investigation identified contextual and high-impact factors that influence ownership choices by altering parent node states and monitoring changes in posterior probabilities. By showing how legal ambiguity spreads throughout the network, this enhanced interpretability. The validation and sensitivity analysis findings are collected in Table 6. During cross-validation, the model achieved more than 80% agreement between inferred and actual judicial outcomes. Robustness testing verified trustworthy inference under limited-evidence settings, and low posterior variance across folds suggested stable, generalizable conditional probabilities. Sensitivity analysis revealed that while AI autonomy and the legality of training data greatly increased attribution uncertainty, contractual ownership clauses and human creative involvement had the greatest impact on ownership attribution. Technical and contractual conditions interacted with jurisdictional variables, exerting a moderate effect. Overall, the findings show that the suggested paradigm is both legally interpretable and empirically sound.

Validation / Sensitivity AspectMethod AppliedMetric / Observation
Case-Based ValidationLeave-one-out cross-validation81.3% ownership outcome consistency
Cross-Fold Stability5-fold cross-validationPosterior variance < 0.05 across folds
Missing Evidence RobustnessPartial evidence inference< 7% change in posterior probabilities
Sensitivity: Human Creative InvolvementOne-factor perturbationΔP(Y) ≈ +0.32
Sensitivity: AI AutonomyOne-factor perturbationΔP(Y) ≈ −0.28
Sensitivity: Training Data LegalityOne-factor perturbationΔP(Y) ≈ +0.25
Sensitivity: Contractual ClausesOne-factor perturbationΔP(Y) ≈ +0.37
Sensitivity: JurisdictionOne-factor perturbationΔP(Y) ≈ +0.14

Table 6: Validation and Sensitivity Analysis Results. Results of model validation and sensitivity analysis, evaluating robustness, stability, and the impact of key variables on ownership inference.

Interpretation and policy implication analysis
The probabilistic outputs of the Bayesian Network provide an interpretable framework for analyzing ownership of AI-generated content by modeling legal and technical factors, including human contribution, AI autonomy, contractual clarity, and the legality of the training data. Rather than serving solely as predictive outputs, the posterior probability estimates support transparent, scenario-based reasoning under uncertain legal conditions. The framework can assist legal professionals in improving consistency during ownership evaluation, support policymakers in identifying regulatory gaps related to AI-generated content, and help AI developers assess how design and governance choices may influence ownership attribution in different legal contexts.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The results of the experiment demonstrate the effectiveness of the Bayesian Network framework in capturing the nature of uncertainty in AI-produced content ownership. The model provides a structured, data-driven approach for supporting ownership inference under uncertainty. The final dataset consisted of 131 cases: a base AI litigation dataset, a public court-opinion database, and an AI copyright case-tracking repository. These cases were selected after applying inclusion, exclusion, and deduplication procedures. The fra...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The proposed Bayesian Network framework constitutes an advancement over current research on ownership of content generated by AI, as opposed to the descriptive or question-identification approaches of other works. The other works, including the base paper by Raghupathi et al. in1, essentially use machine learning approaches in text analytics in attempts to discover themes or categories in the context of AI litigation. Although these approaches enable a clear understanding of what issues of law eme...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no conflicts of interest.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The author acknowledges the academic support and research environment provided by the Faculty of Law, Monash University, which contributed to the completion of this study.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Judicial Opinions & Court Records DatasetCaselaw Access Project (CAP)RRID:SCR_016043Source of court decisions used for extracting legal reasoning and ownership outcomes
AI & Copyright Case TrackerCMS LegalN/ACurated dataset of global AI-related intellectual property disputes
Legal Contracts & Agreements DatasetVarious Public Legal RepositoriesN/ADocuments defining ownership, licensing, and contractual rights
Statutory and Case Law ReferencesGovernment Legal PortalsN/AIntellectual property laws and judicial precedents used for analysis
Legal Annotation GuidelinesExpert-defined (Manual Framework)N/ARules for consistent annotation of legal uncertainty variables
Annotated Legal DatasetGenerated in StudyN/ADataset labeled with legal and technical variables for modeling
Python Programming Environment (v3.10)Python Software FoundationRRID:SCR_008394Core programming environment used for implementation
pgmpy Library (v0.1.24)Open-sourceRRID:SCR_016274Library used for Bayesian Network modeling
NumPy (v1.24)Open-sourceRRID:SCR_008633Numerical computations and array operations
Pandas (v1.5)Open-sourceRRID:SCR_018214Data preprocessing and dataset handling
Jupyter NotebookProject JupyterRRID:SCR_018315Interactive environment for model development
Bayesian Network ModelDeveloped in StudyN/AProbabilistic graphical model for ownership inference
Hill Climbing Algorithm (BIC)Implemented via pgmpyN/AStructure learning method for Bayesian Network
Variable Elimination Inference EngineImplemented via pgmpyN/AComputes posterior ownership probabilities
Scenario Analysis ModuleDeveloped in StudyN/AEnables “what-if” simulations
Sensitivity Analysis ModuleDeveloped in StudyN/AMeasures variable influence on ownership outcomes
Cross-Validation Framework (k-fold, LOOCV)Scikit-learn compatibleRRID:SCR_002577Used for model validation and robustness testing
Computing System (Intel Core i7, 16GB RAM)Intel / OEMN/AHardware used for model training and evaluation
Operating System (Windows 11)Microsoft WindowsRRID:SCR_018096OS environment for implementation

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

AI Content OwnershipOwnership AttributionProbabilistic ReasoningDataset CurationSensitivity AnalysisAI LitigationDecision Support

Related Articles