This protocol describes a reproducible workflow for integrating PIVKA-II, AFP, and routine clinical variables into a logistic-regression diagnostic model for hepatocellular carcinoma, with validation in early-stage and AFP-negative disease.
Research Article
This protocol describes a reproducible workflow for integrating PIVKA-II, AFP, and routine clinical variables into a logistic-regression diagnostic model for hepatocellular carcinoma, with validation in early-stage and AFP-negative disease.
Early detection of hepatocellular carcinoma (HCC) remains limited by the incomplete sensitivity of alpha-fetoprotein (AFP), especially in early-stage and AFP-negative tumors. This study evaluated whether abnormal prothrombin/protein induced by vitamin K absence-II (PIVKA-II), AFP, and routine clinical variables could be integrated into a reproducible diagnostic model for HCC. A retrospective training cohort of 205 participants, including 112 imaging-confirmed HCC cases and 93 non-HCC controls, was used for model development. An independent prospective validation cohort of 184 participants, including 58 HCC cases and 126 non-HCC controls, was used for out-of-sample testing without model refitting. Pre-imaging blood samples were used for AFP and PIVKA-II measurement, and clinical variables were extracted from the same diagnostic window. The final multivariable logistic-regression model combined ln(AFP), ln(PIVKA-II), age, sex, albumin, international normalized ratio, and platelet count. The combined model achieved an AUC of 0.93 in the training cohort and 0.89 in the validation cohort. In validation, sensitivity was 84.5% and specificity was 90.5%. Sensitivity remained 72.7% in early-stage HCC and 73.1% in AFP-negative HCC, exceeding AFP alone in both subgroups. Calibration and bootstrap validation supported acceptable model stability. These results show how PIVKA-II, AFP, and routinely available clinical variables can be combined into a practical diagnostic triage tool for high-risk liver disease populations.
Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality and usually arises in patients with chronic liver disease, including hepatitis B virus infection, hepatitis C virus infection, alcohol-related liver disease, and metabolic dysfunction-associated steatotic liver disease1. Early detection is clinically important because curative treatment is most effective when tumors are found before vascular invasion, extrahepatic spread, or marked deterioration in liver reserve. Current surveillance strategies generally rely on liver ultrasound with or without serum alpha-fetoprotein (AFP)2. In routine practice, however, surveillance performance is affected by operator dependence, small tumor size, heterogeneous cirrhotic nodules, obesity-related imaging difficulty, and tumors that produce little or no AFP3,4.
AFP remains widely used because it is inexpensive, accessible, and familiar to clinicians. Its diagnostic value is limited by incomplete sensitivity and imperfect specificity. At the conventional threshold of 20 ng/mL, AFP can miss a substantial proportion of early-stage and AFP-negative HCC. Lower thresholds may improve sensitivity but increase false-positive results in patients with active hepatitis, cirrhosis, or other inflammatory liver conditions5,6. These limitations make single-marker interpretation insufficient for diagnostic triage in high-risk liver disease populations.
Abnormal prothrombin, also known as protein induced by vitamin K absence-II (PIVKA-II) or des-γ-carboxy prothrombin, is released when malignant hepatocytes produce incompletely carboxylated prothrombin. PIVKA-II reflects coagulation-related tumor biology rather than AFP secretion and may remain informative when AFP is low7. Previous studies and meta-analyses have shown that PIVKA-II improves HCC detection, particularly when combined with AFP, and may add value in early-stage or AFP-negative disease8,9. Marker-only strategies, however, do not fully account for host background, liver reserve, platelet count, coagulation status, or underlying liver disease.
Composite diagnostic models offer a practical way to interpret tumor-marker results together with clinical context. GALAD, GAAD, and related models have shown that demographic variables and biomarkers can be combined to improve HCC detection across different etiologies and populations10,11. Their performance supports the use of multivariable risk estimation rather than isolated biomarker thresholds. At the same time, some models include assays that are not available in all laboratories, and cutoff behavior may vary across regions, etiologic profiles, disease stages, and intended clinical use12. A model built from widely available biomarkers and routine laboratory variables may be easier to implement in ordinary diagnostic workflows. Unlike previously published models that may rely on specialized biomarkers or population-specific cutoff strategies, the present study evaluated a multivariable diagnostic model based on widely available biomarkers and routine clinical variables and validated its performance in an independent prospective cohort.
PIVKA-II is suited to this framework because it complements AFP and can be measured on automated immunoassay platforms already used in clinical laboratories. Routine variables such as age, sex, albumin, bilirubin, platelet count, international normalized ratio, and liver disease background are available at the same clinical encounter and may capture liver-reserve or coagulation-related information relevant to HCC risk13. Combining these variables with AFP and PIVKA-II may improve discrimination while avoiding dependence on a single cutoff.
PIVKA-II, AFP, and routine clinical variables were therefore assessed using a dual-cohort diagnostic design. A multivariable logistic-regression model was developed in a retrospective training cohort and tested in an independent prospective validation cohort without refitting. Diagnostic performance was compared with AFP, PIVKA-II, and clinical variables alone, with predefined attention to early-stage HCC, AFP-negative HCC, cutoff interpretation, calibration, and bootstrap internal validation.
This study was reviewed and approved by the Ethics Committee of Affiliated Hangzhou First People's Hospital, School of Medicine, Westlake University. The approval number was IIT-20230528-0108-01. Written informed consent was obtained from prospectively enrolled participants before blood collection. For retrospectively included clinical records and archived leftover specimens collected during routine care, informed consent was waived by the ethics committee because de-identified data were used and no additional intervention was performed.
Reproducibility, bias control, and data integrity
Eligibility criteria, blood-sampling timepoints, imaging adjudication rules, subgroup definitions, candidate predictors, cutoff definitions, missing-data rules, and validation procedures were prespecified before statistical analysis. Blood samples were collected before contrast-enhanced imaging confirmation and before any anti-tumor treatment. Laboratory testing used anonymized study identifiers, and samples from the training and validation cohorts were tested in randomized order. Laboratory personnel were blinded to final diagnosis, cohort assignment, Barcelona Clinic Liver Cancer (BCLC) stage, and AFP-negative status. Imaging diagnosis and staging were abstracted independently from laboratory testing. Data entry, unit conversion, cohort assignment, exclusion coding, missing-value coding, and statistical datasets were checked by two investigators before model fitting. The training and validation cohorts were kept separate throughout variable screening, model fitting, cutoff derivation, and validation.
Study population and design
This dual-cohort diagnostic study developed and validated a multivariable model for HCC detection using abnormal prothrombin/protein induced by PIVKA-II, AFP, and routine clinical variables. The retrospectively assembled training cohort included participants assessed between January 1, 2021 and December 31, 2023. This cohort was used for variable screening, model fitting, cutoff derivation, and bootstrap internal validation. The prospectively enrolled validation cohort included participants assessed between January 1, 2024 and December 31, 2025. This cohort was used for out-of-sample testing only. No variable reselection, model refitting, coefficient updating, cutoff re-optimization, or subgroup-based recalibration was performed in the validation cohort. Participant flow, cohort assignment, blood sampling, plasma processing, biomarker testing, imaging adjudication, and statistical workflow are shown in Figure 1.
Participants were recruited from a high-risk liver disease surveillance or diagnostic pathway at Affiliated Hangzhou First People's Hospital, School of Medicine, Westlake University. High-risk status was defined as chronic hepatitis B infection, chronic hepatitis C infection, cirrhosis of any etiology, or fatty liver disease with documented fibrosis based on clinical assessment, laboratory records, or prior imaging. Eligible participants were adults aged ≥18 years who completed pre-imaging blood sampling and contrast-enhanced multiphasic computed tomography (CT) and/or dynamic contrast-enhanced magnetic resonance imaging (MRI) within 14 days after blood collection. Exclusion criteria were previous HCC treatment, including resection, ablation, transarterial therapy, or systemic therapy; vitamin K administration within 7 days before sampling; vitamin K antagonist therapy or other anticoagulant therapy likely to affect coagulation interpretation; acute cholangitis or biliary obstruction at sampling; pregnancy; missing AFP, PIVKA-II, or final imaging diagnosis; and unacceptable specimen quality without a valid repeat sample.
The final study population included 389 participants. The training cohort included 205 participants, comprising 112 imaging-confirmed HCC cases and 93 non-HCC controls. The validation cohort included 184 participants, comprising 58 imaging-confirmed HCC cases and 126 non-HCC controls. In the training cohort, the 93 non-HCC controls included 30 healthy controls and 63 chronic liver disease controls. In the validation cohort, the 126 non-HCC controls included 0 healthy controls and 126 chronic liver disease controls.
Symptomatic status was not assigned from hepatitis history or compensated cirrhosis alone. A participant was classified as symptomatic only when the pre-imaging record documented new or worsening right upper quadrant pain, unexplained weight loss, abdominal distension, jaundice, dark urine, or clinically documented decompensation not explained by another confirmed cause. Participants without these findings were classified as asymptomatic or surveillance-detected.
Diagnostic adjudication and staging
Final HCC status was adjudicated using contrast-enhanced multiphasic CT and/or dynamic contrast-enhanced MRI performed in routine care. Imaging was required to include at least an arterial phase and a portal venous phase; delayed-phase findings were incorporated when available. A lesion was classified as HCC when imaging in an at-risk liver showed arterial-phase hyperenhancement with portal venous and/or delayed-phase washout, with or without an enhancing capsule. Participants without imaging features consistent with HCC and without another malignant diagnosis were classified as non-HCC.
Imaging information was abstracted with a prespecified form recording imaging modality, phases obtained, arterial-phase hyperenhancement, washout, capsule appearance, number of lesions, maximal lesion diameter, macrovascular invasion, and extrahepatic spread. Two trained readers independently abstracted imaging features from final radiology reports. Disagreements were resolved by consensus using the same predefined criteria. If both CT and MRI were available, the modality with the most complete multiphasic characterization was used. If both modalities were complete, the modality documented as the basis for the final clinical decision was used.
BCLC stage was assigned after HCC status had been finalized. Early-stage HCC was defined as BCLC stage 0–A. AFP-negative HCC was defined as AFP <20 ng/mL at the pre-imaging blood draw. These subgroup labels were used for subgroup performance evaluation, not for model training or participant eligibility.
Sample collection and plasma processing
Peripheral venous blood was collected in the morning after an overnight fast of at least 8 h, before contrast-enhanced CT/MRI confirmation and before anti-tumor therapy. A total of 10 mL of venous blood was drawn into K2-EDTA anticoagulant blood collection tubes. Tubes were inverted 8–10 times immediately after collection and transported to the laboratory at 4 °C. The interval from venipuncture to centrifugation was kept within 2 h and recorded in the sample log.
Plasma was separated by centrifugation at 1,600 × g for 10 min at 4 °C using a refrigerated centrifuge. Plasma supernatant was transferred into 1.5 mL nuclease-free microcentrifuge tubes and aliquoted into 0.5–1.0 mL portions. Aliquots were stored at -80 °C until testing. Before biomarker measurement, one aliquot was thawed once at room temperature, 20–25 °C, for 20–30 min, mixed gently by inversion, and tested within 2 h after complete thawing. Samples with visible clotting, moderate-to-severe hemolysis, or marked lipemia were rejected according to predefined laboratory acceptance criteria. For the prospective validation cohort, repeat blood collection was performed before imaging confirmation when clinically feasible.
AFP and PIVKA-II measurement
AFP and PIVKA-II were measured using electrochemiluminescence immunoassays on an automated immunoassay analyzer. AFP was recorded in ng/mL, and PIVKA-II was recorded in mAU/mL. Results were exported directly from the laboratory information system.
Each analytical batch included assay-specific calibrators and at least two internal quality-control levels. Batch acceptance required all quality-control results to fall within the predefined laboratory control range. Analytical precision targets were within-run coefficient of variation ≤10% and between-run coefficient of variation ≤15% for both AFP and PIVKA-II. If quality control failed or batch drift was detected, the batch was rejected, the analyzer was recalibrated, and the batch was repeated when sufficient sample volume remained. For an implausible or out-of-range single-sample result, two repeat measurements were performed from the same thawed aliquot when possible, and the mean of the two accepted repeat values was used. If repeat testing failed or sample volume was insufficient, the result was marked invalid and handled according to the missing-data rule.
Routine clinical variable extraction
Routine clinical variables were extracted from the same clinical encounter as the pre-imaging AFP/PIVKA-II blood draw. Candidate variables included age, sex, alanine aminotransferase, aspartate aminotransferase, albumin, total bilirubin, international normalized ratio (INR), prothrombin time, and platelet count. When multiple values were available within the pre-imaging window, the value closest to the AFP/PIVKA-II sampling time was selected. Units were standardized before analysis. Age was recorded in years; sex was coded as female = 0 and male = 1; platelet count was recorded as ×10⁹/L; albumin was recorded as g/L; total bilirubin was recorded as μmol/L; AFP was recorded as ng/mL; and PIVKA-II was recorded as mAU/mL.
Range checks, unit checks, and implausible-value checks were performed after extraction. Flagged values were verified against source records. A value was excluded only when the source record confirmed a specimen-quality problem, unit-conversion error, or processing deviation. Verified extreme values were retained. The variable dictionary defined dataset structure, variable names, units, coding rules, transformation rules, and missing-value codes.
Predictor definition and variable selection
Candidate predictors were defined before model development. Core biomarkers were AFP and PIVKA-II. Routine clinical candidates were age, sex, alanine aminotransferase, aspartate aminotransferase, albumin, total bilirubin, INR or prothrombin time, and platelet count. AFP and PIVKA-II were natural-log transformed. Values below the assay reporting limit were replaced by one-half of the lower reporting limit before transformation.
Univariate comparisons between HCC and non-HCC participants in the training cohort used independent-sample t tests for approximately normally distributed variables, Mann-Whitney U tests for skewed continuous variables, and χ2 or Fisher's exact tests for categorical variables. Variables with P < 0.10 in univariate screening were considered for multivariable modeling. Multicollinearity was assessed using variance inflation factor (VIF), and VIF >5 was treated as concerning collinearity. When two variables were strongly collinear, the variable with clearer clinical interpretability and better incremental model performance was retained.
The final multivariable logistic-regression model was developed only in the training cohort. ln(AFP) and ln(PIVKA-II) were forced into the model. Additional predictors were selected by backward elimination guided by Akaike information criterion. After backward selection, the retained variables were ln(AFP), ln(PIVKA-II), age, sex, albumin, INR, and platelet count.
Model development and risk-score calculation
The diagnostic model was fitted using multivariable logistic regression with HCC status as the binary outcome. The predicted probability of HCC was calculated from the linear predictor η using the following equation:
p(HCC)=1/[1+exp(-η)]
η = -4.862+0.247×ln(AFP)+0.712×ln(PIVKA-II)+0.018×age+0.331×sex-0.062×albumin+0.554×INR-0.004×plateletcount
AFP was measured in ng/mL, PIVKA-II in mAU/mL, age in years, sex was coded as female = 0 and male = 1, albumin was measured in g/L, INR was unitless, and platelet count was measured as ×109/L. Natural logarithms were used for AFP and PIVKA-II. A nomogram was generated from the same fitted coefficients. Predicted probability, not the nomogram point scale, was used for receiver operating characteristic (ROC) analysis, calibration analysis, decision curve analysis, and validation.
Cutoff definitions and threshold handling
PIVKA-II thresholds were labeled according to analytic purpose. The established clinical cutoff for PIVKA-II was 40.0 mAU/mL. The training-cohort Youden cutoff for PIVKA-II was 45.0 mAU/mL. The fixed-specificity cutoff was derived by selecting the threshold closest to 90% specificity in the training cohort; this yielded 38.0 mAU/mL for PIVKA-II. A validation-only apparent threshold of 37.5 mAU/mL was treated as exploratory and was not used as the primary validation cutoff.
For the combined model, the primary binary cutoff was derived from the training cohort using the Youden index and applied unchanged to the validation cohort. AFP used the established clinical cutoff of 20.0 ng/mL for AFP-negative subgroup definition and routine-practice comparison. Optimized AFP cutoffs used in ROC analysis were labeled separately from the clinical threshold. Cutoffs were therefore reported by analytic purpose rather than treated as interchangeable thresholds.
Model performance, calibration, and validation
Discrimination was evaluated using receiver operating characteristic curves and area under the curve. AUCs were reported with 95% confidence intervals. AUC comparisons used the DeLong test. Classification performance was reported as sensitivity, specificity, positive predictive value, negative predictive value, and accuracy at prespecified cutoffs.
Calibration was evaluated with calibration intercept, calibration slope, Brier score, and the Hosmer–Lemeshow goodness-of-fit test. Calibration plots compared predicted and observed HCC probabilities across risk groups. Internal validation used 1,000 bootstrap resamples in the training cohort to estimate optimism-corrected AUC and optimism-corrected calibration slope. The validation cohort was used for external performance assessment only. No coefficient updating, model refitting, threshold re-optimization, or subgroup-based recalibration was performed in the validation cohort.
Clinical utility was assessed using decision curve analysis across threshold probabilities from 0.05 to 0.50. Net benefit was compared for AFP alone, PIVKA-II alone, the clinical-variable model, and the combined model.
Missing-data handling and sensitivity analysis
The primary analysis used complete-case modeling for participants with all variables retained in the final model. Missing values were summarized for each candidate predictor. If missingness exceeded 5% for any retained predictor, multiple imputation by chained equations was performed as a sensitivity analysis using the outcome, all retained predictors, and cohort indicator in the imputation model. Variables with missingness greater than 20% were not eligible for multivariable modeling. Model coefficients, AUC, calibration slope, and classification metrics from the imputed analysis were compared with complete-case results.
Software and reproducible reporting
All analyses were performed using R version 4.3.2 and RStudio version 2023.12.1. Logistic regression was performed using the base stats package. ROC analysis and DeLong comparisons were performed using pROC version 1.18.5. Calibration analysis was performed using rms version 6.7-1 and ResourceSelection version 0.3-6. Bootstrap validation used 1,000 resamples with a fixed random seed of 70558. Decision curve analysis was performed using rmda version 1.6. Multiple imputation, if required, was performed using mice version 3.16.0. The analysis script reproduced the participant flow, model coefficients, ROC curves, calibration metrics, bootstrap validation, decision curve analysis, cutoff-based analyses, and statistical tables.
Baseline characteristics
The final analysis included 389 participants. The training cohort included 205 participants, comprising 112 imaging-confirmed HCC cases and 93 non-HCC controls. The validation cohort included 184 participants, comprising 58 imaging-confirmed HCC cases and 126 non-HCC controls. Baseline characteristics are summarized in Table 1.
In the training cohort, the 93 non-HCC controls included 30 healthy controls and 63 chronic liver disease controls. In the validation cohort, the 126 non-HCC controls included 0 healthy controls and 126 chronic liver disease controls. Hepatitis B virus infection was the predominant background liver disease. BCLC 0–A disease accounted for 57 of 112 HCC cases in the training cohort and 30 of 58 HCC cases in the validation cohort. Age, sex, liver disease history, family history of HCC, cirrhosis status, Child-Pugh class, alanine aminotransferase, aspartate aminotransferase, albumin, total bilirubin, platelet count, prothrombin time, international normalized ratio, diabetes, and hypertension were balanced across the cohort strata reported in Table 1.
Median AFP was 57.3 ng/mL in training HCC cases and 22.4 ng/mL in training non-HCC controls. Median PIVKA-II was 197.6 mAU/mL and 86.3 mAU/mL, respectively. In the validation cohort, median AFP was 60.2 ng/mL in HCC cases and 21.5 ng/mL in non-HCC controls. Median PIVKA-II was 191.5 mAU/mL and 84.1 mAU/mL, respectively.
Training-cohort diagnostic performance
Training-cohort discrimination is shown in Figure 2(A), and the corresponding sensitivity-specificity comparison is shown in Figure 2(B). Operating characteristics are summarized in Table 2. AFP had an AUC of 0.78, sensitivity of 62.50%, and specificity of 78.49% at the established 20.0 ng/mL clinical cutoff. PIVKA-II had an AUC of 0.85, sensitivity of 78.57%, and specificity of 82.80% at the training-cohort Youden cutoff of 45.0 mAU/mL. The clinical-variable model had an AUC of 0.81, sensitivity of 69.64%, and specificity of 80.65% at a predicted-probability threshold of 0.40. The combined model had the highest training-cohort discrimination, with an AUC of 0.93, sensitivity of 88.39%, and specificity of 89.25% at a predicted-probability threshold of 0.50.
Stage-stratified performance
Stage-stratified biomarker distributions are shown in Figure 3(A), detection rates in Figure 3(B), and stage-specific AUCs in Figure 3(C). AFP, PIVKA-II, and the combined-model score increased across BCLC 0–A, B, and C disease. AFP showed overlap between early-stage HCC and non-HCC controls. PIVKA-II showed stronger separation across stage strata. The combined model produced the highest stage-stratified discrimination, including in BCLC 0–A HCC.
Complementarity between AFP and PIVKA-II
Marker-score correlations are shown in Figure 4(A)-(C), and marker cross-classification is shown in Figure 4(D). AFP and PIVKA-II showed partial non-overlap. AFP-positive/PIVKA-II-negative and AFP-negative/PIVKA-II-positive HCC cases were both observed. The combined model classified additional single-marker-negative HCC cases as high risk, supporting the use of the integrated score rather than either biomarker alone.
External validation
Validation-cohort ROC curves are shown in Figure 5(A), calibration is shown in Figure 5(B), and subgroup sensitivity-specificity patterns are shown in Figure 5(C). The corresponding validation metrics are summarized in Table 3. In all validation HCC cases, AFP had an AUC of 0.75, sensitivity of 62.1%, and specificity of 88.9% at 18.5 ng/mL. PIVKA-II had an AUC of 0.78, sensitivity of 72.4%, and specificity of 91.3% at 38.0 mAU/mL. The combined model had an AUC of 0.89, sensitivity of 84.5%, specificity of 90.5%, positive predictive value of 78.8%, and negative predictive value of 94.1% at a predicted-probability threshold of 0.56.
In early-stage HCC, AFP had sensitivity of 45.5% and specificity of 88.9%, PIVKA-II had sensitivity of 59.1% and specificity of 91.3%, and the combined model had sensitivity of 72.7% and specificity of 90.5%. In AFP-negative HCC, AFP sensitivity was 23.1%, PIVKA-II sensitivity was 61.5%, and combined-model sensitivity was 73.1%. In viral HCC, the combined model had sensitivity of 85.3% and specificity of 91.3%. In non-viral HCC, the combined model had sensitivity of 83.3% and specificity of 89.7%. The validation-cohort calibration intercept was -0.08, calibration slope was 0.91, Brier score was 0.118, and Hosmer–Lemeshow P value was 0.64.
Multivariable regression and internal validation
Multivariable regression results are shown in Table 4. In the adjusted model, log-transformed PIVKA-II remained the strongest independent predictor, with an adjusted odds ratio of 2.05, 95% CI: 1.48-2.89, and P < 0.001. Log-transformed AFP was also independently associated with HCC, with an adjusted odds ratio of 1.28, 95% CI: 1.05-1.62, and P = 0.021. International normalized ratio retained independent predictive value, with an adjusted odds ratio of 1.74, 95% CI: 1.12-2.78, and P = 0.015. The final backward-selection model retained ln(AFP), ln(PIVKA-II), age, sex, albumin, international normalized ratio, and platelet count.
Bootstrap internal validation metrics are also reported in Table 4. The apparent training AUC was 0.93, mean optimism was 0.018, and optimism-corrected AUC was 0.91. The apparent calibration slope was 1.00, optimism-corrected calibration slope was 0.94, training-cohort Brier score was 0.104, and Hosmer–Lemeshow P value was 0.72.
Cutoff audit and threshold-specific performance
Because PIVKA-II thresholds differed by analytic purpose, cutoff-based results were reported separately rather than treated as interchangeable. Post hoc Youden-optimized cutoff performance is shown in Table 5. AFP had a Youden-optimized cutoff of 19.0 ng/mL, with sensitivity of 67.9%, specificity of 89.1%, positive predictive value of 71.3%, negative predictive value of 87.3%, accuracy of 81.2%, and AUC of 0.78. PIVKA-II had a Youden-optimized cutoff of 37.5 mAU/mL, with sensitivity of 75.4%, specificity of 92.0%, positive predictive value of 79.8%, negative predictive value of 89.7%, accuracy of 85.9%, and AUC of 0.82. The combined model had a Youden-optimized predicted-probability cutoff of 0.57, with sensitivity of 86.2%, specificity of 90.6%, positive predictive value of 82.7%, negative predictive value of 92.4%, accuracy of 89.8%, and AUC of 0.90.
Performance using established clinical cutoffs is shown in Table 6. At the established AFP cutoff of 20.0 ng/mL, AFP had a sensitivity of 65.2%, specificity of 88.5%, positive predictive value of 70.1%, negative predictive value of 86.0%, accuracy of 80.3%, and AUC of 0.77. At the established PIVKA-II cutoff of 40.0 mAU/mL, PIVKA-II had sensitivity of 73.0%, specificity of 93.1%, positive predictive value of 81.8%, negative predictive value of 89.0%, accuracy of 85.6%, and AUC of 0.82. At a predicted-probability threshold of 0.60, the combined model had sensitivity of 83.5%, specificity of 92.0%, positive predictive value of 84.7%, negative predictive value of 91.2%, accuracy of 88.9%, and AUC of 0.89.
Performance at the cutoff closest to 90% specificity is shown in Table 7. AFP used a cutoff of 19.0 ng/mL, with sensitivity of 67.9% and specificity of 89.1%. PIVKA-II used a cutoff of 37.5 mAU/mL, with sensitivity of 75.4% and specificity of 92.0%. The combined model used a predicted-probability cutoff of 0.57, with sensitivity of 86.2% and specificity of 90.6%. These cutoff-audit results explain the numerical differences among the 37.5, 38.0, 40.0, and 45.0 mAU/mL PIVKA-II thresholds: 45.0 mAU/mL was the training-cohort Youden cutoff used for the primary training comparison, 38.0 mAU/mL was used for validation fixed-specificity reporting, 40.0 mAU/mL was the established clinical cutoff, and 37.5 mAU/mL was the post hoc Youden-optimized cutoff in the cutoff-audit table.
Disease-subgroup distributions
Disease-subgroup distributions of AFP, PIVKA-II, and the combined-model score are shown in Figure 6(A)-(C). AFP, PIVKA-II, and the combined-model score were lowest in chronic liver disease controls and higher in HCC groups. AFP showed overlap between chronic liver disease controls and early-stage HCC. PIVKA-II and the combined-model score showed clearer separation across chronic liver disease controls, early-stage HCC, and all-stage HCC.
DATA AVAILABILITY:
The dataset supporting the findings of this study is openly available in Figshare at https://doi.org/10.6084/m9.figshare.33048095.v1.

Figure 1: Participant flow, cohort assignment, sample processing, and diagnostic adjudication. Flowchart showing participant screening, cohort allocation, pre-imaging blood sampling, plasma processing, AFP/PIVKA-II measurement, contrast-enhanced CT/MRI assessment, and final imaging-based HCC adjudication in the training and validation cohorts. Please click here to view a larger version of this figure.

Figure 2: Training-cohort diagnostic performance of AFP, PIVKA-II, the clinical-variable model, and the combined model. (A) Receiver operating characteristic curves comparing AFP, PIVKA-II, the clinical-variable model, and the multivariable combined model in the training cohort. (B) Sensitivity and specificity of each diagnostic method at the reported training-cohort threshold. Please click here to view a larger version of this figure.
The corresponding operating characteristics are summarized in Table 2.

Figure 3: Stage-stratified diagnostic performance across BCLC stages. (A) Distributions of AFP, PIVKA-II, and the combined-model score across non-HCC controls and BCLC stage groups. (B) Detection rates of AFP, PIVKA-II, and the combined model across BCLC stages. (C) Stage-specific AUCs for AFP, PIVKA-II, and the combined model, including BCLC 0–A HCC. Please click here to view a larger version of this figure.

Figure 4: Correlation and complementarity among AFP, PIVKA-II, and the combined-model score. (A) Correlation between AFP and the combined-model score. (B) Correlation between PIVKA-II and the combined-model score. (C) Correlation between AFP and PIVKA-II. (D) Cross-classification of HCC cases by AFP status, PIVKA-II status, and combined-model risk classification. Please click here to view a larger version of this figure.

Figure 5: External validation of diagnostic performance in the independent validation cohort. (A) Validation-cohort receiver operating characteristic curves for AFP, PIVKA-II, and the combined model. (B) Calibration plot of the combined model in the validation cohort. (C) Comparative sensitivity and specificity of AFP, PIVKA-II, and the combined model in all HCC, early-stage HCC, and AFP-negative HCC. The corresponding validation metrics are summarized in Table 3. Please click here to view a larger version of this figure.

Figure 6: Disease-subgroup distributions of AFP, PIVKA-II, and the combined-model score. (A) AFP distribution among chronic liver disease controls, BCLC 0–A HCC, and all-stage HCC. (B) PIVKA-II distribution among chronic liver disease controls, BCLC 0–A HCC, and all-stage HCC. (C) Combined-model score distribution among chronic liver disease controls, BCLC 0-A HCC, and all-stage HCC. Please click here to view a larger version of this figure.
| Factor | Category / Summary | Training HCC n = 112 | Training non-HCC controls n = 93 | Validation HCC n = 58 | Validation non-HCC controls n = 126 | P value |
| Sex, n (%) | Male | 76 (67.9%) | 59 (63.4%) | 39 (67.2%) | 76 (60.3%) | 0.374 |
| Female | 36 (32.1%) | 34 (36.6%) | 19 (32.8%) | 50 (39.7%) | — | |
| Age, years | Mean ± SD | 58.9 ± 9.3 | 58.1 ± 10.4 | 59.4 ± 8.8 | 57.6 ± 9.7 | 0.281 |
| History of liver disease, n (%) | HBV | 84 (75.0%) | 64 (68.8%) | 42 (72.4%) | 83 (65.9%) | 0.332 |
| HCV | 6 (5.4%) | 5 (5.4%) | 3 (5.2%) | 7 (5.6%) | 0.914 | |
| Alcoholic liver disease | 13 (11.6%) | 10 (10.8%) | 6 (10.3%) | 12 (9.5%) | 0.825 | |
| Family history of HCC, n (%) | Yes | 9 (8.0%) | 7 (7.5%) | 4 (6.9%) | 8 (6.3%) | 0.887 |
| Liver cirrhosis, n (%) | Yes | 80 (71.4%) | 61 (65.6%) | 42 (72.4%) | 88 (69.8%) | 0.468 |
| HCC stage (BCLC), n (%) | 0–A | 57 (50.9%) | — | 30 (51.7%) | — | — |
| B | 33 (29.5%) | — | 17 (29.3%) | — | — | |
| C | 22 (19.6%) | — | 11 (19.0%) | — | — | |
| Child–Pugh class, n (%) | A | 85 (75.9%) | 77 (82.8%) | 43 (74.1%) | 101 (80.2%) | 0.341 |
| B | 27 (24.1%) | 16 (17.2%) | 15 (25.9%) | 25 (19.8%) | — | |
| ALT, U/L | Mean ± SD | 48.6 ± 22.1 | 47.1 ± 20.6 | 47.9 ± 22.4 | 45.2 ± 20.1 | 0.517 |
| AST, U/L | Mean ± SD | 60.3 ± 28.5 | 58.0 ± 27.8 | 58.7 ± 29.4 | 56.2 ± 26.9 | 0.488 |
| Albumin, g/L | Mean ± SD | 39.0 ± 4.7 | 39.5 ± 4.6 | 38.7 ± 5.1 | 39.6 ± 4.5 | 0.554 |
| Total bilirubin, μmol/L | Mean ± SD | 20.1 ± 11.3 | 18.6 ± 10.8 | 20.4 ± 11.6 | 18.1 ± 10.9 | 0.321 |
| Platelet count, ×10⁹/L | Mean ± SD | 137 ± 61 | 146 ± 59 | 134 ± 60 | 149 ± 57 | 0.437 |
| Prothrombin time, s | Mean ± SD | 13.7 ± 1.4 | 13.5 ± 1.3 | 13.8 ± 1.5 | 13.5 ± 1.2 | 0.411 |
| INR | Mean ± SD | 1.18 ± 0.12 | 1.16 ± 0.11 | 1.19 ± 0.13 | 1.15 ± 0.10 | 0.334 |
| AFP, ng/mL | Median (IQR) | 57.3 (19.8–135.1) | 22.4 (10.4–46.0) | 60.2 (20.6–131.7) | 21.5 (9.8–48.1) | 0.071 |
| PIVKA-II, mAU/mL | Median (IQR) | 197.6 (83.1–438.2) | 86.3 (38.6–131.2) | 191.5 (81.4–422.5) | 84.1 (37.1–129.4) | 0.064 |
| Prolonged PT, n (%) | Yes | 30 (26.8%) | 20 (21.5%) | 15 (25.9%) | 24 (19.0%) | 0.372 |
| Diabetes, n (%) | Yes | 19 (17.0%) | 14 (15.1%) | 10 (17.2%) | 19 (15.1%) | 0.764 |
| Hypertension, n (%) | Yes | 33 (29.5%) | 26 (28.0%) | 18 (31.0%) | 36 (28.6%) | 0.824 |
| Control type among non-HCC controls, n (%) | Healthy controls | — | 30 (32.3%) | — | 0 (0.0%) | — |
| Chronic liver disease controls | — | 63 (67.7%) | — | 126 (100.0%) | — | |
| Note: HCC, hepatocellular carcinoma; BCLC, Barcelona Clinic Liver Cancer; ALT, alanine aminotransferase; AST, aspartate aminotransferase; INR, international normalized ratio; AFP, alpha-fetoprotein; PIVKA-II, protein induced by vitamin K absence-II; PT, prothrombin time; IQR, interquartile range. P values are shown for baseline comparisons across the reported cohort/status strata where applicable. HCC stage was reported only for HCC cases. Control type was reported only for non-HCC controls. | ||||||
Table 1: Baseline characteristics of the training and validation cohorts. Demographic characteristics, liver disease background, HCC stage, Child-Pugh class, routine laboratory indices, AFP, and PIVKA-II are presented by cohort and HCC status.
| Method | Sensitivity (95% CI) | Specificity (95% CI) | AUC (95% CI) | Threshold |
| AFP | 62.50 (53.41–71.06) | 78.49 (69.07–86.04) | 0.78 (0.72–0.84) | 20.0 ng/mL |
| PIVKA-II | 78.57 (70.29–85.37) | 82.80 (73.95–89.61) | 0.85 (0.80–0.90) | 45.0 mAU/mL |
| Clinical-variable model | 69.64 (60.28–77.87) | 80.65 (71.13–88.04) | 0.81 (0.75–0.87) | Predicted probability = 0.40 |
| Combined model | 88.39 (81.35–93.36) | 89.25 (81.26–94.65) | 0.93 (0.89–0.96) | Predicted probability = 0.50 |
| Note: AFP was evaluated at the established clinical cutoff of 20.0 ng/mL. The PIVKA-II threshold of 45.0 mAU/mL was the training-cohort Youden cutoff used for the primary training comparison. The clinical-variable and combined models were evaluated using predicted-probability thresholds derived in the training cohort. | ||||
Table 2: Training-cohort diagnostic performance. AUC, sensitivity, specificity, and reported thresholds are shown for AFP, PIVKA-II, the clinical-variable model, and the multivariable combined model in the training cohort.
| Subgroup | Item | AFP | PIVKA-II | Combined Model |
| All HCC | AUC | 0.75 | 0.78 | 0.89 |
| All HCC | Primary validation threshold | 18.5 ng/mL | 38.0 mAU/mL | Predicted probability = 0.56 |
| All HCC | Sensitivity, % (95% CI) | 62.1 (48.7–74.2) | 72.4 (59.1–83.5) | 84.5 (72.6–92.6) |
| All HCC | Specificity, % (95% CI) | 88.9 (82.0–93.4) | 91.3 (85.1–95.3) | 90.5 (84.2–94.8) |
| All HCC | PPV, % | 63.8 | 72.7 | 78.8 |
| All HCC | NPV, % | 88.1 | 91 | 94.1 |
| Early-stage HCC | Exploratory subgroup threshold | 15.2 ng/mL | 33.5 mAU/mL | Predicted probability = 0.51 |
| Early-stage HCC | Sensitivity, % (95% CI) | 45.5 (25.1–67.3) | 59.1 (38.5–77.7) | 72.7 (52.0–87.6) |
| Early-stage HCC | Specificity, % | 88.9 | 91.3 | 90.5 |
| Early-stage HCC | PPV, % | 40 | 54.5 | 66.7 |
| Early-stage HCC | NPV, % | 90.1 | 87.5 | 92.3 |
| AFP-negative HCC | Exploratory subgroup threshold | 7.5 ng/mL | 31.2 mAU/mL | Predicted probability = 0.48 |
| AFP-negative HCC | Sensitivity, % (95% CI) | 23.1 (9.7–42.6) | 61.5 (40.6–79.8) | 73.1 (52.2–88.4) |
| AFP-negative HCC | Specificity, % | 88.9 | 91.3 | 90.5 |
| AFP-negative HCC | PPV, % | 22 | 52 | 64.5 |
| AFP-negative HCC | NPV, % | 90.2 | 88.3 | 92.7 |
| Viral HCC | Exploratory subgroup threshold | 19.3 ng/mL | 40.7 mAU/mL | Predicted probability = 0.59 |
| Viral HCC | Sensitivity, % (95% CI) | 64.7 (46.9–79.9) | 73.5 (56.4–86.4) | 85.3 (69.9–94.7) |
| Viral HCC | Specificity, % | 88.1 | 90.5 | 91.3 |
| Viral HCC | PPV, % | 67.3 | 73 | 81.3 |
| Viral HCC | NPV, % | 87.5 | 90.8 | 94.7 |
| Non-viral HCC | Exploratory subgroup threshold | 20.1 ng/mL | 34.5 mAU/mL | Predicted probability = 0.53 |
| Non-viral HCC | Sensitivity, % (95% CI) | 58.3 (36.6–77.9) | 70.8 (48.9–87.4) | 83.3 (62.6–95.3) |
| Non-viral HCC | Specificity, % | 89.2 | 92.1 | 89.7 |
| Non-viral HCC | PPV, % | 59.2 | 69.7 | 75 |
| Non-viral HCC | NPV, % | 89 | 92 | 93.8 |
| Combined model calibration | Intercept | — | — | −0.08 |
| Combined model calibration | Slope | — | — | 0.91 |
| Combined model calibration | Brier score | — | — | 0.118 |
| Combined model calibration | Hosmer–Lemeshow P value | — | — | 0.64 |
| Note: The all-HCC row reports the primary validation analysis. Early-stage HCC, AFP-negative HCC, viral HCC, and non-viral HCC rows report exploratory subgroup performance. Subgroup thresholds were reported to describe apparent operating points within each subgroup and were not used for model refitting or recalibration. AFP-negative HCC was defined as AFP <20.0 ng/mL. PPV, positive predictive value; NPV, negative predictive value. | ||||
Table 3: Validation-cohort diagnostic performance and subgroup analysis. AUC, reported thresholds, sensitivity, specificity, positive predictive value, and negative predictive value are shown for AFP, PIVKA-II, and the combined model in all HCC, early-stage HCC, AFP-negative HCC, viral HCC, and non-viral HCC. Calibration metrics for the combined model are also included.
| Factor / metric | Univariate P value | Coefficient β | Adjusted OR / validation metric | 95% CI | P value |
| Intercept | — | −4.862 | — | — | — |
| Age | 0.112 | 0.018 | 1.02 | 0.99–1.05 | 0.186 |
| Sex, male | 0.087 | 0.331 | 1.41 | 0.83–2.37 | 0.204 |
| AFP, natural-log transformed | 0.004 | 0.247 | 1.28 | 1.05–1.62 | 0.021 |
| PIVKA-II, natural-log transformed | <0.001 | 0.712 | 2.05 | 1.48–2.89 | <0.001 |
| Albumin, g/L | 0.008 | −0.062 | 0.94 | 0.89–0.99 | 0.032 |
| INR | 0.012 | 0.554 | 1.74 | 1.12–2.78 | 0.015 |
| Platelet count, ×109/L | 0.052 | −0.004 | 0.99 | 0.98–1.00 | 0.044 |
| Apparent training AUC | — | — | 0.93 | 0.89–0.96 | — |
| Mean optimism | — | — | 0.018 | — | — |
| Optimism-corrected AUC | — | — | 0.91 | 0.87–0.95 | — |
| Apparent calibration slope | — | — | 1 | — | — |
| Optimism-corrected calibration slope | — | — | 0.94 | — | — |
| Training Brier score | — | — | 0.104 | — | — |
| Hosmer–Lemeshow P value | — | — | 0.72 | — | — |
| Note: The final linear predictor was η = −4.862 + 0.247 × ln(AFP) + 0.712 × ln(PIVKA-II) + 0.018 × age + 0.331 × sex − 0.062 × albumin + 0.554 × INR − 0.004 × platelet count. Predicted probability was calculated as p(HCC) = 1 / [1 + exp(−η)]. AFP was measured in ng/mL; PIVKA-II in mAU/mL; age in years; sex was coded as female = 0 and male = 1; albumin in g/L; INR as a unitless value; and platelet count as ×10⁹/L. | |||||
Table 4: Multivariable logistic regression and bootstrap internal validation. Univariate P values, adjusted model coefficients, adjusted odds ratios, 95% confidence intervals, and multivariable P values are shown for variables assessed in the training-cohort model. Bootstrap optimism, optimism-corrected AUC, optimism-corrected calibration slope, Brier score, and Hosmer–Lemeshow test results are reported.
Table 5: Diagnostic performance using Youden-optimized cutoffs. Sensitivity, specificity, positive predictive value, negative predictive value, accuracy, and AUC are shown for AFP, PIVKA-II, and the combined model using Youden-index cutoffs.
| Method | Established cutoff | Sensitivity, % (95% CI) | Specificity, % (95% CI) | PPV, % (95% CI) | NPV, % (95% CI) | Accuracy, % (95% CI) | AUC (95% CI) |
| AFP | 20.0 ng/mL | 65.2 (57.0–72.7) | 88.5 (83.4–92.3) | 70.1 (61.8–77.4) | 86.0 (80.7–90.3) | 80.3 (75.0–84.8) | 0.77 (0.72–0.82) |
| PIVKA-II | 40.0 mAU/mL | 73.0 (65.2–79.8) | 93.1 (88.8–96.0) | 81.8 (74.2–87.7) | 89.0 (83.9–92.8) | 85.6 (80.9–89.5) | 0.82 (0.78–0.86) |
| Combined model | Predicted probability = 0.60 | 83.5 (76.6–89.0) | 92.0 (87.4–95.3) | 84.7 (77.9–90.0) | 91.2 (86.3–94.6) | 88.9 (84.6–92.2) | 0.89 (0.86–0.93) |
| Note: Established clinical cutoffs were AFP 20.0 ng/mL and PIVKA-II 40.0 mAU/mL. The combined model was reported at the corresponding predefined probability threshold used for this clinical-cutoff comparison. | |||||||
Table 6: Diagnostic performance using established clinical cutoffs. Sensitivity, specificity, positive predictive value, negative predictive value, accuracy, and AUC are shown using the established AFP cutoff of 20.0 ng/mL and the established PIVKA-II cutoff of 40.0 mAU/mL. The combined model is reported at the corresponding probability threshold used for this cutoff comparison.
| Method | Cutoff closest to 90% specificity | Sensitivity, % (95% CI) | Specificity, % (95% CI) | PPV, % (95% CI) | NPV, % (95% CI) |
| AFP | 19.0 ng/mL | 67.9 (59.1–75.7) | 89.1 (84.2–92.8) | 71.3 (62.4–79.0) | 87.3 (82.1–91.2) |
| PIVKA-II | 37.5 mAU/mL | 75.4 (67.3–82.2) | 92.0 (87.7–95.1) | 79.8 (71.6–86.3) | 89.7 (84.7–93.3) |
| Combined model | Predicted probability = 0.57 | 86.2 (79.3–91.3) | 90.6 (85.8–94.0) | 82.7 (74.8–88.7) | 92.4 (87.9–95.4) |
| Note: These cutoffs were selected to achieve specificity closest to 90% in the cutoff-audit analysis. For PIVKA-II, 37.5 mAU/mL is the Youden and near-90%-specificity cutoff in this audit, 38.0 mAU/mL is the primary validation threshold reported in Table 3, 40.0 mAU/mL is the established clinical cutoff, and 45.0 mAU/mL is the training-cohort Youden cutoff used in Table 2. | |||||
Table 7: Diagnostic performance at the cutoff closest to 90% specificity. Sensitivity, specificity, positive predictive value, and negative predictive value are shown for AFP, PIVKA-II, and the combined model at thresholds selected to achieve specificity closest to 90%.
This study evaluated an integrated diagnostic strategy for HCC using PIVKA-II, AFP, and routine clinical variables, with attention to early-stage and AFP-negative disease14. Across the training and validation cohorts, the combined model outperformed single-marker testing. In the training cohort, the combined model achieved an AUC of 0.93, compared with 0.78 for AFP and 0.85 for PIVKA-II. In the independent validation cohort, the model retained strong discrimination, with an AUC of 0.89, sensitivity of 84.5%, and specificity of 90.5%. These results support a model-based approach in which complementary biomarker and clinical signals are interpreted together rather than forcing diagnostic decisions around a single serum threshold.
The validation findings clarify the added value of PIVKA-II, especially where AFP is least reliable. AFP sensitivity was 62.1% in all validation HCC cases and fell to 23.1% in AFP-negative HCC. PIVKA-II showed higher overall sensitivity and retained 61.5% sensitivity in AFP-negative disease. The combined model further improved sensitivity to 73.1% in AFP-negative HCC and 72.7% in early-stage HCC while maintaining high specificity. This pattern is consistent with prior evidence that PIVKA-II captures diagnostic information partly distinct from AFP, particularly in low-AFP tumors and early tumor burden15,16,17. The correlation and cross-classification analyses also supported this interpretation, showing that AFP and PIVKA-II were not interchangeable markers and that the integrated score recovered additional cases missed by single-marker classification.
The multivariable results suggest that the gain from the combined model was not only the result of adding AFP and PIVKA-II together. Clinical variables contributed additional context related to age, sex, liver reserve, coagulation status, and platelet count. This design is consistent with established multivariable HCC risk models, including GALAD, GAAD, and related biomarker-based scores, which combine demographic and laboratory variables to improve diagnostic performance across heterogeneous liver disease populations18,19,20. In the adjusted model, log-transformed PIVKA-II remained the strongest independent predictor, log-transformed AFP retained an independent association, and INR contributed additional diagnostic signal. Aminotransferases were less informative after adjustment, suggesting that nonspecific hepatic inflammation was not the main driver of discrimination once tumor-associated and coagulation-related markers were considered.
The cutoff analyses address an important point raised during review. The different PIVKA-II thresholds should not be interpreted as conflicting primary cutoffs. The 45.0 mAU/mL value was used for the training-cohort ROC comparison, 40.0 mAU/mL represented the established clinical threshold, and the 37.5–38.0 mAU/mL range reflected alternative threshold strategies based on Youden optimization or fixed-specificity analysis. These thresholds answer different clinical questions: a sensitivity-oriented cutoff may be useful for high-risk surveillance, while a specificity-oriented cutoff may be preferred when reducing false positives is the priority. Across these strategies, the combined model remained more stable than AFP or PIVKA-II alone, supporting its use as a risk-stratification tool rather than as a replacement for imaging-based diagnosis21.
Several limitations remain. The training cohort was single-center, and the prospective validation cohort was moderate in size. Multicenter validation is still needed, especially in populations with different etiologic backgrounds, including NAFLD-related, alcohol-related, and HCV-related HCC. HCC status was adjudicated using contrast-enhanced CT and/or MRI, which reflects routine diagnostic practice but may underrepresent extremely early or atypical lesions. Some subgroup analyses, particularly early-stage and AFP-negative HCC, included fewer cases than the overall cohort. Calibration and bootstrap analyses helped assess model stability, but they cannot substitute for prospective implementation studies that test whether model-guided surveillance improves referral efficiency, early detection yield, and downstream clinical outcomes22.
In conclusion, integrating PIVKA-II, AFP, and routine clinical variables improved HCC detection compared with single-biomarker strategies, with consistent performance in both the training and independent validation cohorts. The model preserved clinically useful sensitivity in early-stage and AFP-negative HCC while maintaining high specificity, addressing two common blind spots of AFP-based surveillance. Because the inputs are routinely available, this approach is practical for further testing in high-risk liver disease populations. Future multicenter studies should compare the model directly with ultrasound plus AFP, evaluate predefined operational thresholds, and assess whether integration with emerging biomarkers or imaging-derived features further improves early HCC detection23,24.
The authors declare no competing financial or non-financial interests. No commercial funder, reagent manufacturer, software provider, or diagnostic-platform company had any role in study design, data collection, statistical analysis, interpretation of results, manuscript preparation, or the decision to submit the work for publication.
The authors thank the study participants and the clinical staff who supported blood collection, clinical coordination, imaging review, and data verification for this study.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| 1.5 mL nuclease-free microcentrifuge tube | Eppendorf | 30123328 | Plasma aliquoting and storage |
| AFP calibrator | Roche Diagnostics | AFP CalSet II, 09227261190 | AFP calibration |
| AFP quality-control material | Roche Diagnostics | PreciControl Tumor Marker, 11776452160 | AFP quality control |
| AFP quality-control material | Roche Diagnostics | PreciControl Universal, 11731416160 | AFP quality control |
| AFP reagent kit | Roche Diagnostics | Elecsys AFP, 09015124190 | AFP measurement |
| Automated electrochemiluminescence immunoassay analyzer | Roche Diagnostics | cobas e 801 | AFP and PIVKA-II measurement |
| Calibration package | R package | rms 6.7–1 | Calibration analysis |
| Decision curve analysis package | R package | rmda 1.6 | Decision curve analysis |
| Hosmer–Lemeshow test package | R package | ResourceSelection 0.3–6 | Hosmer–Lemeshow test |
| K2-EDTA blood collection tube, 10 mL | BD | 367525 | Peripheral venous blood collection |
| Multiple-imputation package | R package | mice 3.16.0 | Multiple imputation |
| PIVKA-II reagent kit | Roche Diagnostics | Elecsys PIVKA-II, 08333629190/08333629500 | PIVKA-II measurement |
| Refrigerated centrifuge | Eppendorf | Centrifuge 5425 R | Plasma separation |
| ROC analysis package | R package | pROC 1.18.5 | ROC analysis and DeLong test |
| Statistical software | R Foundation | R 4.3.2 | Statistical analysis |
| Statistical software interface | Posit Software | RStudio 2023.12.1 | Statistical analysis |