Skip to main navigation Skip to main content

Clin Mol Hepatol : Clinical and Molecular Hepatology

OPEN ACCESS
ABOUT
BROWSE ARTICLES
FOR CONTRIBUTORS

Articles

Original Article

GAFAD: A liquid chromatography-tandem mass spectrometry-based model for early hepatocellular carcinoma detection beyond GALAD’s limitations

Clinical and Molecular Hepatology 2026;32(3):1225-1239.
Published online: February 25, 2026

1R&D Center for Clinical Mass Spectrometry, Seegene Medical Foundation, Seoul, Korea

2Department of Gastroenterology, Ajou University School of Medicine, Suwon, Korea

Corresponding author: Je-Hyun Baek, R&D Center for Clinical Mass Spectrometry, Seegene Medical Foundation, 288 Dapsimni-ro, Dongdaemun-gu, Seoul 02637, Korea Tel: +82-2-2218-9016, Fax: +82-2-2024-4634, E-mail: jhbaek@mf.seegene.com

Editor: Terry Cheuk-Fung Yip, Prince of Wales Hospital, Hong Kong

• Received: October 31, 2025   • Revised: February 8, 2026   • Accepted: February 23, 2026

Copyright © 2026 by Korean Association for the Study of the Liver

This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/3.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

  • 2,206 Views
  • 160 Download
  • 1 Crossref
  • 1 Scopus
prev next
  • Background/Aims
    The GALAD (Gender, Age, Lens culinaris agglutinin-reactive alpha-fetoprotein [AFP-L3], alpha-fetoprotein [AFP], and des-γ-carboxy prothrombin) score, widely used for hepatocellular carcinoma (HCC) detection, was primarily derived from cohorts with advanced-stage tumors and elevated biomarker levels, potentially overestimating accuracy in early-stage disease. Furthermore, the lectin-based AFP-L3 assay has poor sensitivity at low AFP concentrations, limiting detection of small or AFP-negative tumors.
  • Methods
    We developed GAFAD, a multivariable model replacing AFP-L3 with fucosylated AFP percentage, quantified by a validated liquid chromatography–tandem mass spectrometry assay. The model was trained and tested using a hepatitis B virus (HBV)-related cohort (HCC n=235; non-HCC n=290), a diagnostically challenging set with substantial overlap in biomarker levels between HCC and non-HCC. Moreover, a final model (GAFAD) was validated in two independent cohorts (HCC n=210; non-HCC n=245), comprising HBV-, HCV-related and non-viral etiologies.
  • Results
    In the development cohort, GAFAD showed superior diagnostic performance to GALAD for distinguishing HCC from non-HCC, with a higher area under the receiver operating characteristic curve (AUC, 0.938 vs. 0.887; P<0.0001) and greater sensitivity (82% vs. 66%) and accuracy (86% vs. 79%) at 90% specificity. In the external validation cohort, GAFAD similarly outperformed GALAD, achieving a higher AUC (0.874 vs. 0.841, P<0.05), greater sensitivity (72% vs. 57%), and improved accuracy (82% vs. 75%) at 90% specificity. This superiority extended to early-stage, very-early-stage, and AFP-negative HCC.
  • Conclusions
    GAFAD provides a reliable and generalizable tool for early HCC detection across diverse etiologies, supporting its clinical applicability in surveillance and diagnosis.
• GAFAD replaces lectin-based AFP-L3 with mass spectrometry-quantified AFP-Fuc% for HCC detection.
• Validated across diverse etiologies, GAFAD outperforms GALAD and ASAP in viral HCC and shows comparable performance in non-viral disease.
• The model demonstrates superior sensitivity for detecting early-stage and AFP-negative tumors.
• GAFAD offers robust calibration and generalizability, supporting its utility in real-world clinical surveillance.
Graphical Abstract
Hepatocellular carcinoma (HCC) remains a major cause of cancer-related mortality worldwide, particularly in regions with a high prevalence of chronic hepatitis B virus (HBV) infection [13]. Early detection is critical for improving survival, but current surveillance using ultrasound and alpha-fetoprotein (AFP) lacks adequate sensitivity, missing a substantial proportion of early-stage and AFP-negative HCC cases [46].
Several serum biomarkers, including AFP, Lens culinaris agglutinin-reactive alpha-fetoprotein (AFP-L3), and desgamma-carboxy prothrombin (DCP, also known as protein induced by vitamin K absence or antagonist-II [PIVKA-II]), have been proposed to aid HCC detection [7]. However, AFPL3 detection depends on lectin-affinity methods that have limited reproducibility and accessibility in clinical laboratories [811].
The GALAD model, which integrates AFP, AFP-L3, DCP, age, and gender, has improved HCC detection performance [12,13], but its utility is restricted by AFP-L3 assay limitations and inconsistent performance across etiologies and populations [1416].
Liquid chromatography–tandem mass spectrometry (LCMS/MS) offers enhanced analytical specificity and standardization [17,18], and recent studies have demonstrated the feasibility of quantifying fucosylated AFP (AFP-Fuc%) directly without lectin-based enrichment [11,19,20] (Supplementary Fig. 1).
This study developed and validated GAFAD, a novel MS-based algorithm that replaces lectin-derived AFP-L3 with AFP-Fuc% for HCC detection. The diagnostic performance of GAFAD was evaluated in development and independent validation cohorts across different HCC etiologies and compared with the established GALAD model.
Study population
This study was designed as a retrospective case–control analysis using de-identified serum samples obtained from institutional biobanks. The study included a development cohort and independent external validation cohorts, as summarized in Figure 1. Blood samples were collected within 1 month of HCC diagnosis in 96% of patients (median: –3 days, range: −127 to +18 days). The development cohort consisted of serum samples provided by the Ajou University Hospital Biobank, a member of the Korea Biobank Network. Samples were collected between January 2014 and December 2019 and distributed after anonymization in accordance with institutional ethical and procedural standards. This cohort was used exclusively for model development and internal validation. Inclusion criteria, disease definitions, and diagnostic standards were applied consistently with previously established protocols and are summarized below. Healthy controls were adults aged 18–50 years who underwent routine health checkups and had no known history of liver disease or chronic medical conditions. Chronic HBV infection was defined by the persistence of hepatitis B surface antigen for at least six months. Liver cirrhosis (LC) was diagnosed based on characteristic ultrasonographic findings, liver stiffness measurements (>14 kPa), and supportive laboratory abnormalities, including thrombocytopenia and/or hypoalbuminemia [21]. HCC was diagnosed according to imaging-based or histological criteria in accordance with the Korean Liver Cancer Association guidelines [6]. Tumor stage was determined using the modified Union for International Cancer Control (mUICC) staging system [22,23]. All HCC patients in the development cohort were treatment-naïve at the time of sample collection. To ensure etiologic homogeneity for model training, only HBV-related liver disease and HCC cases were included in the development cohort.
For external validation, independent serum samples were obtained from two separate institutional biobanks (Fig. 1): The Keimyung University Dongsan Hospital Biobank (samples collected between 2013 and 2024) and Ajou University Hospital Biobank (samples collected between 2017 and 2021, independent of those used for model development). The external validation cohorts included patients with chronic liver disease and HCC of diverse etiologies, including HBV, hepatitis C virus (HCV), and non-viral causes. These samples were used solely for external validation of the previously developed models and were not involved in model training or parameter estimation.
Measurements of HCC biomarkers
All biomarker measurements were performed using previously validated analytical methods. Measurements from the development cohort were obtained as part of our earlier study [11] and were reused for the present multi-marker model development. In addition, all samples included in the external validation cohorts were newly processed and analyzed using the same analytical workflow. For the development cohort, a total of 525 serum samples from HBV-related patients at Ajou University Hospital were analyzed, with no exclusions due to missing clinical data or LC-MS/MS assay failure. AFP and AFP-Fuc% were simultaneously quantified using a validated LC-MS/MS assay, as previously described [11]. For GALAD score calculation, AFP and AFP-L3 levels were measured using a lectin-based immunoassay (μTAS i30; Wako, Osaka, Japan), and DCP was measured using a chemiluminescent microparticle immunoassay (MultiSR; Abbott Diagnostics, Abbott Park, IL, USA). Serum samples from the external validation cohorts were processed independently but analyzed using the same LCMS/MS instrumentation, sample preparation procedures, calibration strategy, and quality control criteria as applied to the development cohort. Immunoassay-based measurements of AFP, AFP-L3, and DCP for both the development and external validation cohorts were performed independently at the Seegene Medical Foundation central laboratory, according to the manufacturers’ instructions. In the external validation cohort, AFP-L3 could not be quantified in one HCC patient (mUICC stage I) whose extremely high AFP exceeded the assay’s measurable range even after dilution.
Model development
Data analysis and model development were performed using logistic regression modeling within a machine learning framework implemented in the R software (version 4.3.2, The R Foundation). Model training and evaluation were conducted using the caret package (version 6.0–94). To minimize sampling bias and ensure independent model validation, the dataset was randomly divided into a training set (75%) and a test set (25%) using the createDataPartition function, which applies stratified sampling to ensure proportional representation of the normal, liver disease, and HCC groups. The random seed was fixed to ensure reproducibility. Model training and hyperparameter optimization were conducted exclusively within the training set. To enhance model robustness and generalizability and reduce overfitting, ten-fold cross-validation repeated three times was applied to the training set using the trainControl function. The final optimized models were subsequently evaluated on the independent test set, which was not used during model training or tuning.
For comparison, GALAD scores were calculated using a previously published equation [12]:
Z=1.67×Gender+0.09×Age+0.04×AFP-L3 (%)+2.34×log10 [AFP (ng/mL)]+1.33×log10[DCP (mAU/mL)]−10.08, where gender was coded as 1 for males and 0 for females.
To overcome the limited diagnostic sensitivity of single biomarkers, we developed new multivariable scoring models incorporating multiple biomarkers, including lectin-independent fucosylated AFP measured by LC-MS/MS (AFPFuc%, denoted as F). Two diagnostic scores, GAFA and GAFAD, were constructed using logistic regression. The GAFA model incorporated four variables: gender, age, AFP-Fuc%, and AFP, where AFP-Fuc% was defined as the ratio of fucosylated AFP to total glycosylated AFP. The GAFAD model included DCP as an additional biomarker, resulting in a five-variable model. The final equations, which are protected by patent, were defined as follows:
Equation 1
GAFAD=α1Gender+α2Age+α3AFP-Fuc+α4log10[AFP (ng/mL)]+α5log10[DCP (mAU/mL)]+α6
Equation 2
GAFA=β1Gender+β2Age+β3AFP-Fuc+β4log10[AFP (ng/mL)]+β5
The regression coefficients for each variable included in the GAFA and GAFAD models were estimated using the training set and are summarized in Supplementary Table 1. Coefficient estimates, standard errors, odds ratios, and corresponding p values are provided to describe the direction and magnitude of association between each predictor and HCC risk. These coefficients formed the basis of the final scoring equations and were fixed for subsequent internal and external validation. To enhance interpretability of the multivariable models, feature contribution was further assessed using SHapley Additive exPlanations (SHAP). SHAP values were calculated to quantify the contribution of each predictor to the model output at the individual and population levels. For each variable, the distribution of SHAP values and the mean absolute SHAP value were used to summarize the direction and relative importance of feature effects on HCC prediction (Supplementary Fig. 2). This analysis provided an intuitive assessment of how individual biomarkers and clinical variables influenced the GAFAD model beyond regression coefficients.
Statistical analysis
Diagnostic performance was evaluated by comparing the new LC-MS/MS-based models with the conventional immunoassay-based GALAD score. Because GALAD requires AFP-L3, the one HCC patient without a quantifiable AFP-L3 value was excluded from the between-model comparisons (n=454; 209 HCC vs. 245 non-HCC), whereas cohort-level descriptions used the full set (n=455). The primary endpoint was discrimination performance as assessed by the area under the receiver operating characteristic curve (AUC). Secondary performance metrics included the area under the precision–recall curve (AUC-PR), F1 score, Youden index, calibration plots, integrated discrimination improvement (IDI), and net reclassification improvement (NRI), which were used to provide complementary perspectives on model performance. AUC was used to evaluate overall discrimination between HCC and non-HCC cases, with 95% confidence intervals (CI). AUC-PR and F1 scores were additionally assessed to better reflect model performance under class imbalance conditions, where sensitivity and precision are particularly relevant. Calibration plots were used to assess the agreement between predicted probabilities and observed event rates across risk strata. IDI and NRI were applied to quantify incremental improvements in discrimination and risk classification relative to the GALAD model. These performance metrics were not applied as independent hypothesis-testing procedures but were used as complementary validation measures addressing distinct aspects of diagnostic accuracy, consistent with prior methodological evaluations of biomarker-based prediction models. Formal adjustment for multiple comparisons was therefore not applied. Normality of continuous variables was assessed using the Shapiro–Wilk test, and nonparametric comparisons were performed using the Mann–Whitney U-test. All statistical analyses were conducted using MedCalc (version 19.7) and R software (version 4.3.2). A two-tailed P-value <0.05 was considered statistically significant where applicable.
Ethics statement
This study was conducted in accordance with the ethical principles of the Declaration of Helsinki (1975; revised 2013). The study protocol was reviewed and approved by the Institutional Review Board of Ajou University Hospital, Suwon, South Korea (IRB No. AJOUIRB-KSP-2016-365), which covered the development cohort, and by the Institutional Review Board of Seegene Medical Foundation, Seoul, South Korea (SMF-IRB-2023-002), which covered the independent validation cohorts. The requirement for informed consent was waived owing to the retrospective nature of the study and anonymization of patient data.
Data availability
The data generated and analyzed during this study are available from the corresponding author upon reasonable request. Data sharing is restricted to data types for which a community-recognized, structured public repository does not exist and is subject to institutional and ethical approval, in accordance with biobank and IRB regulations.
Patient characteristics
A total of 525 previously collected datasets were used for model development and internal validation, as described in our prior study of individual HCC biomarkers [11]. The cohort was randomly divided into a training set (n=395, 75.2%) and a test set for internal validation (n=130, 24.8%) (Fig. 1). The training set comprised 36 healthy controls, 182 patients with chronic liver disease, and 177 patients with HCC (Supplementary Table 2). The test set included 14 healthy controls, 58 patients with chronic liver disease, and 58 patients with HCC. For external validation, an independent case–control cohort (n=455) was obtained from two institutional biobanks: Keimyung University Dongsan Hospital Biobank and Ajou University Hospital Biobank (Supplementary Table 3). These samples were collected independently from the development cohort and used exclusively for external validation (Fig. 1). Tumor stage was classified according to mUICC staging system. Very earlystage HCC was defined as mUICC stage I (single tumor ≤ 2 cm), a category that broadly overlaps with Barcelona Clinic Liver Cancer (BCLC) stage 0. Early-stage HCC was defined as mUICC stage II; this category broadly overlaps with, but is not equivalent to, BCLC stage A.
Evaluation of individual biomarkers
We first evaluated the diagnostic performance of individual biomarkers to assess their ability to distinguish HCC cases from non-HCC cases. This analysis was performed using both established clinical thresholds and previously optimized cutoffs. The optimal cutoff for AFP-Fuc% was set at 8.3%, based on a threshold determined in our previous study [11]. For AFP and DCP, clinical cutoffs of 20 ng/mL and 40 mAU/mL, respectively, were applied according to the established guidelines [24,25]. As shown in the distribution plots (Supplementary Fig. 3), the median values of the single markers (AFP, AFP-Fuc%, and DCP) were consistently higher in HCC cases than in non-HCC cases. This trend was also observed across tumor stages, with progressively higher marker levels in more advanced stages of HCC. However, single biomarkers demonstrated limited discriminative performance between the HCC and non-HCC groups. This finding is supported by the sensitivity of each marker when applying established clinical cutoffs, which ranged from 47% to 58%, based on AUC analysis.
Diagnostic assessment of new scoring models
The performance of the developed scoring models in distinguishing HCC cases from non-HCC cases was evaluated. The optimal cutoffs were determined using the Youden index and were −0.39 for GAFAD and −0.56 for GAFA. At these thresholds, the sensitivity and specificity of GAFAD were 83% and 89%, respectively, whereas GAFA achieved a sensitivity of 82% and a specificity of 74% (Supplementary Fig. 4A, 4B). Unlike individual biomarkers, the scoring models demonstrated markedly improved discriminatory ability at their respective optimal cutoffs, with the sensitivity increasing by more than 1.5-fold.
When compared with GALAD, the newly developed models, GAFA and GAFAD, outperformed GALAD not only at its established cutoff (−0.63) but also when evaluated using the Youden index–optimized threshold (−0.03) (Supplementary Fig. 4C) [12]. When diagnostic performance was assessed by fixing the specificity at 90%, the sensitivity for each marker and score differed substantially (Supplementary Table 4). AFP showed a sensitivity of 46%, AFP-Fuc% showed 53%, and DCP and GALAD reached 66%. Notably, GAFAD demonstrated the highest sensitivity (82%), indicating superior performance in high-specificity clinical settings. A similar trend was observed for precision and overall accuracy, both of which increased across the markers in the same order. Compared to GALAD, GAFAD achieved a 16% increase in sensitivity, 7% absolute improvement in accuracy, and 3% gain in precision, further supporting the enhanced diagnostic utility of the GAFAD model over existing single and multi-marker approaches.
Comparative diagnostic performance across tumor stage–defined and AFP-stratified subgroups
Comparative analyses were conducted to evaluate the diagnostic performance of the scoring models across tumor stage–defined and AFP-stratified clinical subgroups in both the development cohort (training and internal test sets) and the independent external validation cohorts (Fig. 2). Subgroup analyses included the total cohort (HCC vs. non-HCC), early-stage HCC (mUICC stage II), very-early-stage HCC (mUICC stage I), and AFP-negative HCC (AFP <20 ng/mL), each compared with liver disease controls. In the development cohort, diagnostic performance was first assessed in the total population. Single biomarkers demonstrated modest discriminatory ability, whereas multi-marker models achieved higher AUCs, with GAFAD showing the best performance, followed by GALAD and GAFA (Table 1, first column). Similar performance trends were observed in the internal test set, confirming model robustness. When analyses were restricted to tumor stage–defined subgroups, multi-marker models consistently outperformed single biomarkers in both early-stage and very early-stage HCC. In patients with early-stage HCC (mUICC stage II), GAFAD demonstrated superior diagnostic accuracy compared with GAFA and GALAD (Table 1, second column). This advantage was maintained in very earlystage HCC (mUICC stage I), indicating robust performance even in the earliest clinically detectable disease (Table 1, third column).
Given that median AFP levels in most early-stage subgroups remained below the conventional diagnostic cutoff (Supplementary Fig. 3A), we further evaluated diagnostic performance in AFP-negative HCC. In this subgroup, multimarker models retained discriminatory power, whereas single biomarkers showed limited sensitivity. Although DCP exhibited relatively high AUCs in the training set, this finding was not consistently reproduced in the test set. In contrast, GAFAD consistently achieved the highest AUC across both training and test sets (Table 1, fourth column).
These subgroup-specific performance patterns were reproducibly observed in the external validation cohorts. As shown in Figure 2E–2H, GAFAD consistently demonstrated superior diagnostic accuracy across total, early-stage, very-early-stage and AFP-negative HCC subgroups, confirming the generalizability of the model beyond the development cohort. To further assess robustness across independent institutions, we performed site-specific analyses within the external validation cohorts. As shown in Supplementary Figure 5, the diagnostic performance of GAFAD was consistently maintained in both Keimyung University Dongsan Hospital and Ajou University Hospital cohorts, with no meaningful performance degradation across sites. These findings indicate that the observed performance of GAFAD is not driven by a single institution and is reproducible across different clinical settings. Across all clinically relevant subgroups, GAFAD achieved the highest AUC with statistically significant differences, whereas GAFA and GALAD showed comparable performance.
Etiology-specific diagnostic performance
To evaluate diagnostic performance according to disease etiology, analyses were performed by stratifying patients into HBV-related, HCV-related, and non-viral HCC subgroups. Because the development cohort consisted exclusively of HBV-related cases, the primary HBV analysis was conducted using the combined cohort, whereas analyses for HCV-related and non-viral HCC were based solely on the external validation cohorts. In each etiology-specific analysis, HCC cases were compared with etiology-matched liver disease controls, including chronic hepatitis and LC. In HBV-related disease, GAFAD demonstrated the highest diagnostic accuracy, achieving an AUC of 0.923 and outperforming GALAD and other multi-marker models (Fig. 3A). In HCV-related HCC, GAFAD also showed the highest AUC (0.910), although the difference compared with GALAD did not reach statistical significance (Fig. 3B). In non-viral HCC, the diagnostic performance of GAFAD was comparable to that of GALAD, with no statistically significant difference between the two models (Fig. 3C). To further assess model performance under clinically challenging conditions, we performed additional etiology-specific analyses restricted to AFP-negative HCC (AFP <20 ng/mL). In these analyses, all etiology-specific comparisons were conducted using the external validation cohorts only, consistent with the approach used for HCV-related and non-viral disease. As shown in Supplementary Figure 6, GAFAD consistently demonstrated the highest diagnostic performance in viral etiologies, including both HBV and HCV, whereas its performance in non-viral HCC remained comparable to that of GALAD and other composite scores without evidence of inferior discrimination. Overall, GAFAD maintained robust diagnostic performance across viral and non-viral etiologies, including AFP-negative disease, and did not show inferior discrimination compared with established multi-marker models in any etiology-specific subgroup.
Association of GAFAD and individual biomarkers with tumor burden and vascular invasion
To explore the relationship between biomarker levels and tumor characteristics, we evaluated the distributions of GAFAD, AFP, AFP-Fuc%, and PIVKA-II according to vascular invasion status and tumor size in patients with HCC (Supplementary Fig. 7). Comparisons were performed using the Mann–Whitney U-test for vascular invasion and the Kruskal–Wallis test for tumor size categories, with additional assessment of monotonic trends using Spearman correlation analysis. Patients with vascular invasion showed significantly higher GAFAD scores compared with those without vascular invasion (median 11.65 vs. 3.07, P<0.0001). Similar patterns were observed for AFP-Fuc% and PIVKA-II, both of which were significantly elevated in the presence of vascular invasion (P<0.01). In contrast, AFP levels did not differ significantly according to vascular invasion status. When stratified by tumor size, GAFAD demonstrated a strong stepwise increase across increasing tumor size categories (≤2 cm, 2–5 cm, 5–10 cm, and >10 cm; P<0.0001), with a robust positive correlation between GAFAD score and tumor size (Spearman ρ=0.606, P<0.0001). AFP-Fuc% and PIVKA-II also showed significant increases with tumor size and strong positive correlations (both P<0.0001). By contrast, AFP exhibited only a modest association with tumor size, with a weaker correlation coefficient (Spearman ρ=0.197, P<0.01). Overall, among the evaluated biomarkers and composite scores, GAFAD showed the most pronounced dynamic range and strongest association with both vascular invasion and increasing tumor size, suggesting enhanced sensitivity to changes in tumor burden.
Comparison of model performance across complementary statistical metrics
To comprehensively assess diagnostic performance, we compared the multi-marker models GAFAD and GAFA with GALAD and ASAP (age, sex, AFP, PIVKA-II) [26] using complementary statistical metrics, including precision–recall analysis, F1 score, reclassification indices, and calibration measures. Because subgroup analyses involved class imbalance and reduced sample sizes, particularly in earlystage and AFP-negative HCC, model performance was primarily evaluated using the AUC-PR, which more accurately reflects true-positive detection under such conditions.
In the overall population, AUC-PR values increased progressively from GAFA (0.869; 95% CI 0.840–0.897) to GALAD (0.881; 95% CI 0.853–0.907), ASAP (0.920; 95% CI 0.892–0.946), and GAFAD (0.933; 95% CI 0.896–0.952), with GAFAD showing the highest precision across a wide range of recall levels (Supplementary Table 5). Consistent improvements were also observed in F1 scores, indicating a better balance between precision and recall. When stratified into training, test, and external validation sets, GAFAD consistently outperformed GALAD and ASAP in distinguishing HCC from non-HCC cases (Supplementary Fig. 8A). Similar trends were maintained in clinically relevant subgroups, including very-early-stage HCC, early-stage HCC, and AFP-negative HCC compared with liver disease controls (Supplementary Fig. 8B–8D), demonstrating robustness under diagnostically challenging conditions. Reclassification and calibration analyses further supported the superiority of GAFAD.
Compared with GALAD, GAFAD achieved significantly higher IDI=0.506 (P<0.001) and NRI=1.407 (P<0.001), indicating substantial gains in risk discrimination and classification accuracy (Supplementary Table 5). GAFA and ASAP also showed improved performance over GALAD, although to a lesser extent. Calibration analysis demonstrated improved agreement between predicted probabilities and observed outcomes for GAFAD, as reflected by better model fit statistics and a nonsignificant Hosmer–Lemeshow test (P=0.321), whereas GALAD showed evidence of miscalibration (P<0.001). Collectively, these results indicate that GAFAD provides consistently superior discrimination, reclassification, and calibration performance across multiple statistical metrics and clinically relevant subgroups.
Comparison of biomarker distribution characteristics between our cohort and previous GALAD studies
To assess the diagnostic complexity of our study cohort relative to prior GALAD-based studies, we compared the distributions of three key biomarkers (AFP, AFP-L3, and DCP) between the HCC and non-HCC groups across six HBV-related cohorts: our dataset and four published GALAD studies. In our cohort, the median AFP value in HCC patients was 16.3 ng/mL (interquartile range [IQR] 4.7–297.1) compared to 4.5 ng/mL (IQR 2.8–8.9) in non-HCC, representing only a ~3.6-fold difference and substantial IQR overlap. In contrast, previously published studies have reported significantly greater separation between groups. For example, one study reported a median AFP level of 57 ng/mL in HCC versus 2.8 ng/mL in non-HCC, with minimal IQR overlap [14]. Similar patterns were observed for AFP-L3 and DCP; other cohorts exhibited markedly elevated biomarker levels in HCC patients and wider separation between groups [12,13,27,28].
Figure 4 presents a dot plot illustrating AUC values plotted against the summed median differences of the biomarkers (AFP, AFP-L3, and DCP) between the HCC and non-HCC groups. The hyperbolic model (R2=0.972) was well-fitted across various studies, which indicated that GALAD performance increased as biomarker separation widened but declined sharply in cohorts with minimal separation. Our cohort showed the narrowest total median difference (16.9), whereas other studies exhibited substantially larger median differences (2–9 times). These findings confirm that our dataset presents a diagnostically more challenging scenario, characterized by narrower intergroup biomarker differences and greater overlap, especially reflective of early-stage and AFP-negative HCC cases. Despite these constraints, our GAFAD model (open circle) demonstrated a notable improvement in diagnostic performance, with the AUC increasing from 0.887 (GALAD) to 0.938 (ΔAUC=0.051), as indicated by the arrow in Figure 4.
In this study, we developed and validated two LC-MS/ MS–based multivariable models (GAFA and GAFAD) for the early detection of HCC, using AFP-Fuc% as a useful biomarker. Both models outperformed individual markers and the widely used GALAD score, affirming the diagnostic advantage of integrating complementary serological markers. Among them, GAFAD demonstrated the most consistent and clinically relevant improvements, particularly in the very-early-stage and AFP-negative HCC groups, which may be often missed by current surveillance strategies.
GAFAD retains the GALAD framework but introduces two key advancements: substitution of AFP-L3 with AFP-Fuc% and quantification via LC-MS/MS. This change addresses two major limitations of the GALAD models. First, AFP-L3, measured using lectin-based assays, shows poor sensitivity at low AFP levels and limited utility in early HCC [29]. In contrast, AFP-Fuc%, previously validated using LC-MS/MS, provides highly sensitive and reliable measurements across the entire AFP range, including AFP-negative cases [11]. Second, GALAD was developed from cohorts dominated by relatively advanced-stage HCC with highly elevated biomarkers, leading to overestimated performance when applied to early-stage detection [12].
We suggest that GAFAD’s advantage reflects information gain where traditional assays are least informative: the low-AFP zone. LC–MS/MS-based AFP-Fuc% maintains a continuous, low-noise signal below conventional reporting thresholds, enabling the model to exploit fine-grained risk gradients and interactions with AFP and DCP values. Consequently, GAFAD improves discrimination and reclassification at high-specificity operating points that matter for early detection, while maintaining superior calibration— features that collectively can explain its clinical utility beyond the mere reduction of missing values.
Our model was trained in a diagnostically challenging cohort composed predominantly of patients with early-stage AFP-negative HCC. This population exhibited substantial biomarker overlap with non-HCC cases, closely simulating clinical conditions in which accurate early diagnosis is the most difficult. We found a nonlinear fit with high fidelity in the plot of GALAD performance for multiple studies regarding the median gaps between the disease groups. This reveals a strong functional relationship between the median differences in biomarkers and GALAD performance. Interestingly, the studies included sample cohorts from multiple regions, such as America, Europe, Africa, and Asia [12,13,27,28], suggesting a minimal influence of ethnicity. However, the biomarker distribution profiles of the selected cohorts, such as differences in median concentrations and the degree of range overlap between disease groups, are likely to be the most critical factors in evaluating or developing diagnostic score algorithms. Notably, even under these challenging conditions, GAFAD demonstrated superior AUCs compared to GALAD, with gains of +0.051 for overall HCC vs. non-HCC, +0.061 for early-stage HCC, +0.094 for very-early-stage HCC and +0.128 for AFP-negative HCC. Specifically, GAFAD exhibited its most robust predictive power in early-stage HCC, yielding the highest AUC values compared to other diagnostic settings such as very-early-stage or AFP-negative HCC in both development and external validation cohorts. Moreover, the GAFAD maintained consistent cutoff values across subgroups and demonstrated improved calibration, highlighting its robustness and readiness for clinical integration.
Our study highlights that GAFAD exhibits superior diagnostic accuracy compared to current scoring algorithms, including ASAP and GALAD. Specifically, GAFAD showed robust performance in early and very-early-stages, as well as in AFP-negative HCC (Supplementary Fig. 4). In HCC patients with normal-range AFP (<7 ng/mL, n=83) in the development cohort, GAFAD detected more than twice as many cases (46/83, 55.4%) as ASAP (22/83, 26.5%) and substantially more than GALAD (27/83, 32.5%). This exceptional sensitivity implies that GAFAD establishes a new benchmark for early diagnosis, particularly overcoming the limitations of conventional surveillance in patients with non-elevated biomarkers and ambiguous ultrasound observations.
Beyond its diagnostic utility, GAFAD also reflects tumor burden and aggressiveness. As illustrated in Supplementary Figure 7, we observed a progressive increase in GAFAD scores correlating with tumor size and number. This trend is primarily driven by the contribution of AFP-Fuc% and DCP levels. Furthermore, patients with vascular invasion exhibited significantly higher GAFAD scores compared to those without. Collectively, these findings indicate that the GAFAD score holds potential not only for early detection but also as a valuable biomarker for monitoring disease progression and predicting prognosis. This contrasts sharply with conventional single biomarkers, such as AFP alone, which often lack the sensitivity to capture such pathological aggressiveness, highlighting the distinct advantage of GAFAD’s integrative approach in clinical practice.
Despite its development within a homogeneous HBV-related cohort, GAFAD demonstrates promise for broader application. Our external validation across two independent cohorts not only confirmed its efficacy in HBV patients but also supported its generalizability to HCV and non-viral HCC cases. It is important to note, however, that GAFAD’s performance was somewhat attenuated in the non-viral population, a limitation also observed in the GALAD score [13]. This phenomenon may stem from the intrinsic reliance on AFP and DCP within the linear regression framework, indicating that the pathophysiology of metabolic dysfunction–associated steatohepatitis (MASH)/metabolic dysfunction–associated steatotic liver disease (MASLD) driven HCC might require distinct or supplementary biomarkers. To address this, future prospective studies across diverse ethnicities and etiologies would be beneficial to refine the model and confirm its global clinical applicability.
To further validate the model’s robustness against potential clinical confounders, we specifically investigated the influence of hepatic inflammation and antiviral therapy on biomarker levels. Although antiviral therapy was more frequently administered in HCC patients (35.2% vs. 26.1%), stratified analyses indicated that aminotransferase (AST/ALT) levels were comparable or higher in the control liver disease group. Crucially, AFP levels remained significantly elevated in HCC patients regardless of inflammatory status, confirming that AFP elevation in our cohort primarily reflects tumor-related secretion rather than hepatic necroinflammation.
From a translational perspective, GAFAD enables the sensitive detection of HCC even in AFP-negative patients, supports integration into pan-etiology and cirrhosis surveillance workflows, and leverages a fully quantitative, low-missing-value LC-MS/MS platform. While mass spectrometry involves a more complex laboratory infrastructure compared to immunoassays, its growing adoption in clinical laboratories and the potential for centralized testing make implementation increasingly feasible [17,18].
Although the external validation cohorts included patients with non-viral liver disease, including metabolic-associated etiologies, further stratification of MASH-related HCC was limited, particularly among cancer cases, due to the inherent constraints of biobank-based sample annotation. Future prospective studies with etiologically well-defined MASLD/MASH cohorts will be required to clarify performance in this growing population.
In parallel, the GAFA model excluding DCP showed comparable or superior performance in population-level screening, particularly when distinguishing HCC patients from healthy controls (AUC up to 0.991 in certain comparisons) as shown in Supplementary Figure 9. Given its simplified biomarker panel, GAFA may offer a cost-effective option for broad liver cancer screening when DCP testing is unavailable.
In summary, GAFAD was optimized under strictly constrained diagnostic conditions and consistently outperformed established tools. Beyond its superior sensitivity in early detection, GAFAD’s significant correlation with tumor burden and vascular invasion highlights its dual utility as both a diagnostic instrument and a prognostic monitoring marker. Its biologically informed design, robustness across clinically challenging subgroups, and high accuracy support its use as a practical and generalizable solution for improving early HCC detection, monitoring disease progression, and ultimately expanding access to curative therapy.

Authors’ Contributions

H. Kim wrote the original manuscript. J.H. Baek co-drafted and critically reviewed the manuscript. H. Kim and J.H. Baek were responsible for the conception and design of the study, interpretation of the data, and editing of the manuscript. H. Kim set up and operated automated sample preparation. J. Park performed LC-MRM analysis. H. Kim performed the statistical analyses. W. Oh performed the machine-learning analysis. S. Lee, W. S. Yang, S. S. Kim, and J. Y. Cheong have reviewed the manuscript.

Acknowledgements

The biospecimens and data used in this study were provided by the Biobank of Keimyung University Dongsan Medical Center and the Biobank of Ajou University Hospital, which is a member of the Korea Biobank Network. Abstract figure was generated by Nano Banana Pro and English grammar was corrected by ChatGPT.

Conflicts of Interest

The authors have no conflicts to disclose.

Supplementary material is available at Clinical and Molecular Hepatology website (http://www.e-cmh.org).

Supplementary Figure 1.

Schematic illustration of fucosylated α-fetoprotein (AFP-Fuc%) quantification by LC-MS/MS. Non-fucosylated (AFP₀) and fucosylated (AFP₁) glycopeptides are detected, and AFP-Fuc% is calculated as: AFP-Fuc%=[AFP₁/(AFP₀+AFP₁)]×100. The red triangle represents the core α1-6 fucose residue. LC-MS/MS, liquid chromatography–tandem mass spectrometry.
cmh-2025-1244-Supplementary-Fig-1.pdf

Supplementary Figure 2.

SHAP-based interpretation of the GAFAD model. Distribution of SHAP values for each predictor (left) and mean absolute SHAP values indicating relative feature importance (right). SHAP analysis illustrates the contribution and directionality of individual variables to the GAFAD model output. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; SHAP, SHapley Additive exPlanations.
cmh-2025-1244-Supplementary-Fig-2.pdf

Supplementary Figure 3.

Serum levels of alpha-fetoprotein (AFP), fucosylated AFP (AFP-Fuc), and des-γ-carboxy prothrombin (DCP) across clinical groups. (A) AFP (ng/mL), (B) AFP-Fuc (%), and (C) DCP (mAU/mL) concentrations were measured in healthy controls, patients with chronic hepatitis B (HBV), liver cirrhosis (LC), and hepatocellular carcinoma (HCC) stratified by tumor stage (HCC stage1, HCC stage2, and HCC stage3). Data are plotted on a logarithmic scale, with each dot representing an individual sample. Horizontal black bars indicate median values. Red horizontal lines denote clinical cut-off thresholds for each biomarker (AFP: 20 ng/mL; AFP-Fuc: 8.3%; DCP: 40 mAU/mL), along with corresponding sensitivity (Sen.) and specificity (Spe.) values. Statistical significance between groups was assessed using the Mann–Whitney U-test; P-values are indicated.
cmh-2025-1244-Supplementary-Fig-3.pdf

Supplementary Figure 4.

Distribution of multivariable models (GAFAD, GAFA, and GALAD) across clinical groups. (A) GAFAD (B) GAFA and (C) GALAD scores were calculated in healthy controls, patients with chronic hepatitis B (HBV), liver cirrhosis (LC), and hepatocellular carcinoma (HCC) stratified by tumor stage (HCC stage1, HCC stage2, and HCC stage3). Each dot represents an individual sample and horizontal black bars indicate median values. (A, B) Dashed horizontal lines denote GAFAD and GAFA cut-off values (–0.39 and –0.56, respectively), and (C) two dashed horizontal lines denote conventional and Youden index GALAD cut-off values (–0.03 and –0.63), with corresponding sensitivity (Sen.) and specificity (Spe.) values. Statistical significance between groups was assessed using the Mann–Whitney U-test; P-values are indicated. GAFA, gender, age, alpha-fetoprotein (AFP), and fucosylated alpha-fetoprotein (AFP-Fuc); GAFAD, gender, age, AFP, AFP-Fuc, and des-γ-carboxy prothrombin (DCP); GALAD, gender, age, AFP, Lens culinaris agglutinin-reactive alpha-fetoprotein (AFP-L3), and DCP.
cmh-2025-1244-Supplementary-Fig-4.pdf

Supplementary Figure 5.

Site-specific diagnostic performance of individual biomarkers and multi-marker models in the external validation cohort. Receiver operating characteristic curves comparing AFP, AFP-Fuc%, DCP, GAFA, GALAD, and GAFAD in external validation samples from Ajou University Hospital (A) and Keimyung University Dongsan Hospital (B). In both institutional cohorts, GAFAD demonstrated consistently high diagnostic performance, indicating robustness and reproducibility across independent clinical sites. Area under the curve values are shown within each panel. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP.
cmh-2025-1244-Supplementary-Fig-5.pdf

Supplementary Figure 6.

Etiology-specific diagnostic performance in AFP-negative hepatocellular carcinoma (HCC). Receiver operating characteristic curves comparing the diagnostic performance of GALAD, GAFAD, and ASAP in AFP-negative HCC (AFP <20 ng/mL), stratified by disease etiology: HBV-related (A), HCV-related (B), and non-viral HCC (C). All analyses were performed using the external validation cohorts only. In each etiology-specific analysis, AFP-negative HCC cases were compared with etiology-matched liver disease controls, including chronic hepatitis and liver cirrhosis. Area under the curve values are shown within each panel. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; ASAP, age, sex, AFP, PIVKA-II; DCP, des-γ-carboxy prothrombin; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; HBV, hepatitis B virus; HCV, hepatitis C virus; PIVKA-II, protein induced by vitamin K absence or antagonist-II.
cmh-2025-1244-Supplementary-Fig-6.pdf

Supplementary Figure 7.

Association of GAFAD and individual biomarkers with vascular invasion and tumor size in HCC. Box plots showing the distributions of GAFAD scores, AFP, AFP-Fuc%, and PIVKA-II according to vascular invasion status (upper panels) and tumor size categories (lower panels) in patients with hepatocellular carcinoma. For vascular invasion analyses, biomarker levels were compared between patients without vascular invasion (VI−) and those with vascular invasion (VI+). For tumor size analyses, patients were stratified into four groups (≤2 cm, 2–5 cm, 5–10 cm, and >10 cm). Group comparisons were performed using the Mann–Whitney U-test for vascular invasion and the Kruskal–Wallis test for tumor size categories. Spearman correlation coefficients were calculated to assess monotonic associations with tumor size. Median values and sample sizes are indicated in each panel. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; HCC, hepatocellular carcinoma; PIVKA-II, protein induced by vitamin K absence or antagonist-II.
cmh-2025-1244-Supplementary-Fig-7.pdf

Supplementary Figure 8.

Precision–recall analysis of multi-marker models across datasets and clinically relevant subgroups. Precision– recall curves comparing GAFAD, GAFA, GALAD, and ASAP in the training set, internal test set, and external validation set for discrimination between HCC and non-HCC (A). Subgroup analyses include comparisons between very-early-stage HCC and liver disease (B), early-stage HCC and liver disease (C), and AFP-negative HCC (AFP <20 ng/mL) and liver disease (D). Precision (positive predictive value) is plotted against recall (sensitivity). Area under the precision–recall curve values are indicated in each panel. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; ASAP, age, sex, AFP, PIVKA-II; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; HCC, hepatocellular carcinoma; PIVKA-II, protein induced by vitamin K absence or antagonist-II.
cmh-2025-1244-Supplementary-Fig-8.pdf

Supplementary Figure 9.

Receiver operating characteristic curves comparing GAFAD, ASAP, GALAD, and GAFA models for hepatocellular carcinoma (HCC) detection. (A) HCC vs. non-HCC, (B) very-early-stage HCC (single nodule ≤2 cm) vs. liver disease (chronic hepatitis B and liver cirrhosis), and (C) very-early-stage HCC vs. healthy controls. The red line represents the GAFAD model, yellow the ASAP model, black the GALAD model, and blue the GAFA model. Area under the curve (AUC) values for each model are shown in the figure panels. Dashed diagonal lines represent the line of no discrimination (AUC=0.5). Asterisks indicate statistical significance between models based on DeLong’s test (P<0.05). AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutininreactive alpha-fetoprotein; ASAP, age, sex, AFP, PIVKA-II; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; PIVKA-II, protein induced by vitamin K absence or antagonist-II.
cmh-2025-1244-Supplementary-Fig-9.pdf

Supplementary Table 1.

Regression coefficients and odds ratios (OR) for variables included in the GAFAD model
cmh-2025-1244-Supplementary-Table-1.pdf

Supplementary Table 2.

Baseline demographic, clinical, and biomarker characteristics of the study population in the training and test sets
cmh-2025-1244-Supplementary-Table-2.pdf

Supplementary Table 3.

Baseline demographic, clinical, and biomarker characteristics of the external validation cohorts
cmh-2025-1244-Supplementary-Table-3.pdf

Supplementary Table 4.

Diagnostic performance metrics of individual biomarkers and composite models for hepatocellular carcinoma (HCC) detection
cmh-2025-1244-Supplementary-Table-4.pdf

Supplementary Table 5.

Comparative statistical performance of GALAD, GAFAD, and GAFA models for hepatocellular carcinoma (HCC) detection
cmh-2025-1244-Supplementary-Table-5.pdf
Figure 1
Study design and workflow for GAFA(D) score development and validation. The development cohort (n=525), consisting predominantly of HBV-related liver disease and HCC cases from Ajou University Hospital, was randomly divided into a training set (n=395) and an internal test set (n=130). LC–MS/MS–based biomarker analysis and logistic regression modeling were performed in the training set to develop the GAFA(D) score, which was internally validated in the test set. Two independent external validation cohorts (total n=455), from Keimyung University Dongsan Hospital and Ajou University Hospital, were used to evaluate the diagnostic performance of the fixed GAFAD model across diverse etiologies. GAFAD, gender, age, alpha-fetoprotein, fucosylated alpha-fetoprotein, and des-γ-carboxy prothrombin; HBV, hepatitis B virus; HCC, hepatocellular carcinoma; LC, liver cirrhosis; LC–MS/MS, liquid chromatography–tandem mass spectrometry.
cmh-2025-1244f1.jpg
Figure 2
Diagnostic performance of scoring models across tumor stage–defined and AFP-stratified subgroups. Receiver operating characteristic curves comparing individual biomarkers and multi-marker scoring models in the development cohort (A–D) and independent external validation cohorts (E–H). Panels (A and E) show total HCC versus non-HCC; (B and F), early-stage HCC versus liver disease controls; (C and G), very-early-stage HCC versus liver disease controls; and (D and H), AFP-negative HCC versus liver disease controls. Early-stage HCC was defined as mUICC stage II, very-early-stage HCC as mUICC stage I, and AFP-negative HCC as AFP <20 ng/mL. Area under the curve values are shown within each panel, and statistical comparisons between models were performed using the DeLong test. Panels (E and G) exclude one stage I HCC case with a non-quantifiable AFP-L3 value. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; HCC, hepatocellular carcinoma; mUICC, modified Union for International Cancer Control.
cmh-2025-1244f2.jpg
Figure 3
Etiology-specific diagnostic performance of individual biomarkers and multi-marker models. Receiver operating characteristic curves comparing AFP, AFP-Fuc%, DCP, GAFA, GALAD, and GAFAD according to disease etiology. Analyses were performed separately for HBV-related HCC (A), HCV-related HCC (B), and non-viral HCC (C). For HBV-related disease, the cohort was combined with development and external validation cohort, whereas analyses for HCV-related and non-viral HCC were based on the external validation cohorts only. In each etiology-specific analysis, HCC cases were compared with etiology-matched liver disease controls, including chronic hepatitis and liver cirrhosis. Area under the curve values are shown within each panel. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; HBV, hepatitis B virus; HCC, hepatocellular carcinoma; HCV, hepatitis C virus.
cmh-2025-1244f3.jpg
Figure 4
Relationship between biomarker distribution separation and GALAD model performance. Scatter plot showing the association between the sum of median differences (Δ median) in biomarker levels between HCC and non-HCC groups and the area under the curve (AUC) values achieved by the GALAD model across different studies. Data points represent the present study (GALAD and GAFAD) and previously published GALAD studies [12,13,27,28], labeled by model name and reference number. The dashed curve represents the best-fit non-linear regression (R2=0.9717). In the present study, despite a small Δ median (indicating greater overlap in biomarker concentration’s distributions), the GAFAD model (open circle) achieved a higher AUC than GALAD (solid circle), highlighting its robustness in more challenging diagnostic settings. GAFAD, gender, age, alpha-fetoprotein (AFP), fucosylated alpha-fetoprotein (AFP-Fuc), and des-γ-carboxy prothrombin (DCP); GALAD, gender, age, AFP, Lens culinaris agglutinin-reactive alpha-fetoprotein (AFP-L3), and DCP; HCC, hepatocellular carcinoma.
cmh-2025-1244f4.jpg
cmh-2025-1244f5.jpg
Table 1
Diagnostic performance of individual biomarkers and composite models for hepatocellular carcinoma (HCC) in the training and internal test sets
Table 1
(1) HCC vs. non-HCC (2) Early-stage HCC (Stage II) vs. liver disease (3) Very-early-stage HCC (single node, ≤2 cm) vs. liver disease (4) AFP-negative HCC vs. liver disease
Cutoff value AUC (95% CI) Sensitivity (%) Specificity (%) Cutoff value AUC (95% CI) Sensitivity (%) Specificity (%) Cutoff value AUC (95% CI) Sensitivity (%) Specificity (%) Cutoff value AUC (95% CI) Sensitivity (%) Specificity (%)
AFP (established cutoff) 20 0.716 (0.669–0.760) 46.89 88.53 20 0.725 (0.641–0.809) 50.8 86.3 20 0.659 (0.596–0.718) 36.07 86.26 20 0.549 (0.486–0.609)
 AFP-Fuc% (Youden index) 8.3 0.736 (0.690–0.779) 48.59 94.95 8.3 0.733 (0.646–0.819) 49.2 94 8.3 0.616 (0.552–0.677) 32.79 94.51 8.3 0.634 (0.574–0.691) 27.66 94.51
 DCP (established cutoff) 40 0.846 (0.807–0.880) 61.02 98.17 40 0.855 (0.786–0.924) 67.8 97.3 40 0.716 (0.655–0.772) 29.51 97.8 40 0.826 (0.776–0.869) 53.19 97.8
 GALAD (established cutoff) −0.63 0.879 (0.843–0.910) 85.31 66.51 −0.63 0.877 (0.827–0.927) 89.8 59.9 −0.63 0.779 (0.722–0.830) 73.77 60.44 −0.63 0.741 (0.685–0.792) 72.34 60.44
 GALAD (Youden index) −0.03 0.879 (0.843–0.910) 79.66 76.61 0.48 0.877 (0.827–0.927) 76.3 81.3 1.27 0.779 (0.722–0.830) 49.18 92.86 −0.49 0.741 (0.685–0.792) 71.28 63.74
 GAFAD (Youden index) 0.11 0.937 (0.908–0.959) 78.53 95.41 0.64 0.933 (0.891–0.975) 78 97.8 0.11 0.880 (0.832–0.918) 63.93 94.51 0.11 0.882 (0.838–0.918) 64.89 94.51
 GAFA (Youden index) −0.03 0.866 (0.829–0.898) 71.19 83.03 0.26 0.851 (0.793–0.910) 71.2 85.7 −0.79 0.804 (0.748–0.852) 85.25 59.34 −1.2 0.742 (0.686–0.792) 91.49 47.8
Test set
 AFP (established cutoff) 20 0.747 (0.663–0.819) 46.55 87.5 20 0.713 (0.579–0.846) 42.9 82.8 20 0.698 (0.583–0.797) 42.11 84.48 20 0.513 (0.404–0.620)
 AFP-Fuc% (Youden index) 8.3 0.781 (0.700–0.848) 51.72 94.44 8.3 0.853 (0.752–0.954) 57.1 91.4 8.3 0.625 (0.508–0.733) 31.58 93.1 8.3 0.645 (0.537–0.744) 22.58 93.1
 DCP (established cutoff) 40 0.845 (0.772–0.903) 48.28 98.61 40 0.857 (0.746–0.969) 57.1 96.6 40 0.802 (0.696–0.884) 15.79 98.28 40 0.755 (0.653–0.840) 29.03 98.28
 GALAD (established cutoff) −0.63 0.910 (0.847–0.953) 91.38 75 −0.63 0.904 (0.839–0.968) 100 67.2 −0.63 0.803 (0.696–0.885) 73.68 68.97 −0.63 0.801 (0.703–0.878) 83.87 68.97
 GALAD (Youden index) −0.03 0.910 (0.847–0.953) 79.31 83.33 −0.46 0.904 (0.839–0.968) 100 72.4 1.25 0.803 (0.696–0.885) 52.63 94.83 −0.47 0.801 (0.703–0.878) 83.87 72.41
 GAFAD (Youden index) −0.69 0.947 (0.893–0.978) 91.38 88.89 −0.26 0.973 (0.940–1.000) 100 91.4 −0.69 0.888 (0.796–0.948) 84.21 86.21 −0.69 0.894 (0.811–0.950) 83.87 86.21
 GAFA (Youden index) −0.56 0.905 (0.841–0.949) 86.21 79.17 −0.18 0.916 (0.840–0.993) 90.5 82.8 −0.75 0.830 (0.727–0.906) 78.95 70.69 −0.56 0.821 (0.725–0.894) 80.65 74.14

Diagnostic performance was evaluated across multiple clinically relevant conditions, including the total cohort (HCC vs. non-HCC), early-stage HCC (mUICC stage II), very earlystage HCC (mUICC stage I), and AFP-negative HCC (AFP <20 ng/mL). For each biomarker or model, the cutoff value, area under the receiver operating characteristic curve (AUC) with 95% confidence interval (CI), sensitivity, and specificity are shown. “Established cutoff” values represent conventional clinical thresholds, whereas “Youden index” cutoffs were derived by maximizing Youden’s J statistic. Bold values indicate the highest AUC within each condition. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; mUICC, modified Union for International Cancer Control.

AFP

alpha-fetoprotein

AFP-Fuc

fucosylated alpha-fetoprotein

AFP-L3

Lens culinaris agglutinin-reactive alpha-fetoprotein

ASAP

age, sex, AFP, PIVKA-II

AUC

area under the receiver operating characteristic curve

BCLC

Barcelona Clinic Liver Cancer

DCP

des-γ-carboxy prothrombin

GAFA

gender, age, AFP, and AFP-Fuc

GAFAD

gender, age, AFP, AFP-Fuc, and DCP

GALAD

gender, age, AFP, AFP-L3, and DCP

HBV

hepatitis B virus

HCC

hepatocellular carcinoma

IDI

integrated discrimination improvement

IRB

Institutional Review Board

LC

liver cirrhosis

LC–MS/MS

liquid chromatography–tandem mass spectrometry

mUICC

modified Union for International Cancer Control

NRI

net reclassification improvement

PIVKA-II

protein induced by vitamin K absence or antagonist-II

PR

precision–recall
  • 1. Llovet JM, Kelley RK, Villanueva A, Singal AG, Pikarsky E, Roayaie S, et al. Hepatocellular carcinoma. Nat Rev Dis Primers 2021;7:6.
  • 2. Stella L, Santopaolo F, Gasbarrini A, Pompili M, Ponziani FR. Viral hepatitis and hepatocellular carcinoma: from molecular pathways to the role of clinical surveillance and antiviral treatment. World J Gastroenterol 2022;28:2251-2281.
  • 3. Kim BH, Park JW. Epidemiology of liver cancer in South Korea. Clin Mol Hepatol 2018;24:1-9.
  • 4. Singal AG, Zhang E, Narasimman M, Rich NE, Waljee AK, Hoshida Y, et al. HCC surveillance improves early detection, curative treatment receipt, and survival in patients with cirrhosis: a meta-analysis. J Hepatol 2022;77:128-139.
  • 5. Tzartzeva K, Obi J, Rich NE, Parikh ND, Marrero JA, Yopp A, et al. Surveillance imaging and alpha fetoprotein for early detection of hepatocellular carcinoma in patients with cirrhosis: a meta-analysis. Gastroenterology 2018;154:1706-1718e1.
  • 6. Korean Liver Cancer Association (KLCA) and National Cancer Center (NCC) Korea. 2022 KLCA-NCC Korea practice guidelines for the management of hepatocellular carcinoma. J Liver Cancer 2023;23:1-120.
  • 7. Marrero JA, Feng Z, Wang Y, Nguyen MH, Befeler AS, Roberts LR, et al. Alpha-fetoprotein, des-gamma carboxyprothrombin, and lectin-bound alpha-fetoprotein in early hepatocellular carcinoma. Gastroenterology 2009;137:110-118.
  • 8. Noda K, Miyoshi E, Kitada T, Nakahara S, Gao CX, Honke K, et al. The enzymatic basis for the conversion of nonfucosylated to fucosylated alpha-fetoprotein by acyclic retinoid treatment in human hepatoma cells: activation of alpha1–6 fucosyltransferase. Tumour Biol 2002;23:202-211.
  • 9. Li D, Mallory T, Satomura S. AFP-L3: a new generation of tumor marker for hepatocellular carcinoma. Clin Chim Acta 2001;313:15-19.
  • 10. Kudo M. Alpha-fetoprotein-L3: useful or useless for hepatocellular carcinoma? Liver Cancer 2013;2:151-152.
  • 11. Kim H, Park J, Suh H, Lee S, Park Y, Yang WS, et al. Development and validation of a lectin-independent liquid chromatography-tandem mass spectrometry method for serum glycosylated alpha-fetoprotein analysis and comparison with a liquid-phase binding assay. Ann Lab Med 2026;46:62-71.
  • 12. Johnson PJ, Pirrie SJ, Cox TF, Berhane S, Teng M, Palmer D, et al. The detection of hepatocellular carcinoma using a prospectively developed and validated model based on serological biomarkers. Cancer Epidemiol Biomarkers Prev 2014;23:144-153.
  • 13. Yang JD, Addissie BD, Mara KC, Harmsen WS, Dai J, Zhang N, et al. GALAD score for hepatocellular carcinoma detection in comparison with liver ultrasound and proposal of GALADUS score. Cancer Epidemiol Biomarkers Prev 2019;28:531-538.
  • 14. Cagnin S, Donghia R, Martini A, Pesole PL, Coletta S, Shahini E, et al. Galad score as a prognostic marker for patients with hepatocellular carcinoma. Int J Mol Sci 2023;24:16485.
  • 15. Huang CF, Kroeniger K, Wang CW, Jang TY, Yeh ML, Liang PC, et al. Surveillance imaging and GAAD/GALAD scores for detection of hepatocellular carcinoma in patients with chronic hepatitis. J Clin Transl Hepatol 2024;12:907-916.
  • 16. Guan MC, Zhang SY, Ding Q, Li N, Fu TT, Zhang GX, et al. The performance of GALAD score for diagnosing hepatocellular carcinoma in patients with chronic liver diseases: a systematic review and meta-analysis. J Clin Med 2023;12:949.
  • 17. Hua KF, Wu YH, Zhang ST. Clinical diagnostic value of liquid chromatography-tandem mass spectrometry method for primary aldosteronism in patients with hypertension: a systematic review and meta-analysis. Front Endocrinol (Lausanne) 2022;13:1032070.
  • 18. Onwuka DC, Chen LYC, Zhan SH, Seidman MA, Cartagena L, Killow V, et al. Mass spectrometry in IgG4-related disease diagnosis. Sci Rep 2024;14:2584.
  • 19. Kim KH, Lee SY, Hwang H, Lee JY, Ji ES, An HJ, et al. Direct monitoring of fucosylated glycopeptides of alpha-fetoprotein in human serum for early hepatocellular carcinoma by liquid chromatography-tandem mass spectrometry with immunoprecipitation. Proteomics Clin Appl 2018;12:e1800062.
  • 20. Kim KH, Lee SY, Baek JH, Lee SY, Kim JY, Yoo JS. Measuring fucosylated alpha-fetoprotein in hepatocellular carcinoma: a comparison of μTAS and parallel reaction monitoring. Proteomics Clin Appl 2021;15:e2000096.
  • 21. European Association for the Study of the Liver. EASL Clinical Practice Guidelines on non-invasive tests for evaluation of liver disease severity and prognosis - 2021 update. J Hepatol 2021;75:659-689.
  • 22. Ueno S, Tanabe G, Nuruki K, Hamanoue M, Komorizono Y, Oketani M, et al. Prognostic performance of the new classification of primary liver cancer of Japan (4th edition) for patients with hepatocellular carcinoma: a validation analysis. Hepatol Res 2002;24:395-403.
  • 23. Korean Liver Cancer Association, National Cancer Center. 2018 Korean Liver Cancer Association-National Cancer Center Korea Practice Guidelines for the management of hepatocellular carcinoma. Gut Liver 2019;13:227-299.
  • 24. Marrero JA, Kulik LM, Sirlin CB, Zhu AX, Finn RS, Abecassis MM, et al. Diagnosis, staging, and management of hepatocellular carcinoma: 2018 practice guidance by the American Association for the Study of Liver Diseases. Hepatology 2018;68:723-750.
  • 25. Beudeker BJB, Fu S, Balderramo D, Mattos AZ, Carrera E, Diaz J, et al. Validation and optimization of AFP-based biomarker panels for early HCC detection in Latin America and Europe. Hepatol Commun 2023;7:e0264.
  • 26. Yang T, Xing H, Wang G, Wang N, Liu M, Yan C, et al. a novel online calculator based on serum biomarkers to detect hepatocellular carcinoma among patients with hepatitis B. Clin Chem 2019;65:1543-1553.
  • 27. Eletreby R, Elsharkawy M, Taha AA, Hassany M, Abdelazeem A, El-Kassas M, et al. Evaluation of GALAD score in diagnosis and follow-up of hepatocellular carcinoma after local ablative therapy. J Clin Transl Hepatol 2023;11:334-340.
  • 28. Burciu C, Șirli R, Bende R, Popa A, Vuletici D, Miuțescu B, et al. A statistical approach to the diagnosis and prediction of HCC using CK19 and Glypican 3 biomarkers. Diagnostics (Basel) 2023;13:1253.
  • 29. Hou J, Berg T, Vogel A, Piratvisuth T, Trojan J, De Toni EN, et al. Comparative evaluation of multimarker algorithms for earlystage HCC detection in multicenter prospective studies. JHEP Rep 2025;7:101263.

Download Citation

Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

Format:

Include:

GAFAD: A liquid chromatography-tandem mass spectrometry-based model for early hepatocellular carcinoma detection beyond GALAD’s limitations
Clin Mol Hepatol. 2026;32(3):1225-1239.   Published online February 25, 2026
Download Citation

Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

Format:
Include:
GAFAD: A liquid chromatography-tandem mass spectrometry-based model for early hepatocellular carcinoma detection beyond GALAD’s limitations
Clin Mol Hepatol. 2026;32(3):1225-1239.   Published online February 25, 2026
Close

Figure

  • 0
  • 1
  • 2
  • 3
  • 4
GAFAD: A liquid chromatography-tandem mass spectrometry-based model for early hepatocellular carcinoma detection beyond GALAD’s limitations
Image Image Image Image Image
Figure 1 Study design and workflow for GAFA(D) score development and validation. The development cohort (n=525), consisting predominantly of HBV-related liver disease and HCC cases from Ajou University Hospital, was randomly divided into a training set (n=395) and an internal test set (n=130). LC–MS/MS–based biomarker analysis and logistic regression modeling were performed in the training set to develop the GAFA(D) score, which was internally validated in the test set. Two independent external validation cohorts (total n=455), from Keimyung University Dongsan Hospital and Ajou University Hospital, were used to evaluate the diagnostic performance of the fixed GAFAD model across diverse etiologies. GAFAD, gender, age, alpha-fetoprotein, fucosylated alpha-fetoprotein, and des-γ-carboxy prothrombin; HBV, hepatitis B virus; HCC, hepatocellular carcinoma; LC, liver cirrhosis; LC–MS/MS, liquid chromatography–tandem mass spectrometry.
Figure 2 Diagnostic performance of scoring models across tumor stage–defined and AFP-stratified subgroups. Receiver operating characteristic curves comparing individual biomarkers and multi-marker scoring models in the development cohort (A–D) and independent external validation cohorts (E–H). Panels (A and E) show total HCC versus non-HCC; (B and F), early-stage HCC versus liver disease controls; (C and G), very-early-stage HCC versus liver disease controls; and (D and H), AFP-negative HCC versus liver disease controls. Early-stage HCC was defined as mUICC stage II, very-early-stage HCC as mUICC stage I, and AFP-negative HCC as AFP <20 ng/mL. Area under the curve values are shown within each panel, and statistical comparisons between models were performed using the DeLong test. Panels (E and G) exclude one stage I HCC case with a non-quantifiable AFP-L3 value. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; HCC, hepatocellular carcinoma; mUICC, modified Union for International Cancer Control.
Figure 3 Etiology-specific diagnostic performance of individual biomarkers and multi-marker models. Receiver operating characteristic curves comparing AFP, AFP-Fuc%, DCP, GAFA, GALAD, and GAFAD according to disease etiology. Analyses were performed separately for HBV-related HCC (A), HCV-related HCC (B), and non-viral HCC (C). For HBV-related disease, the cohort was combined with development and external validation cohort, whereas analyses for HCV-related and non-viral HCC were based on the external validation cohorts only. In each etiology-specific analysis, HCC cases were compared with etiology-matched liver disease controls, including chronic hepatitis and liver cirrhosis. Area under the curve values are shown within each panel. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; HBV, hepatitis B virus; HCC, hepatocellular carcinoma; HCV, hepatitis C virus.
Figure 4 Relationship between biomarker distribution separation and GALAD model performance. Scatter plot showing the association between the sum of median differences (Δ median) in biomarker levels between HCC and non-HCC groups and the area under the curve (AUC) values achieved by the GALAD model across different studies. Data points represent the present study (GALAD and GAFAD) and previously published GALAD studies [12,13,27,28], labeled by model name and reference number. The dashed curve represents the best-fit non-linear regression (R2=0.9717). In the present study, despite a small Δ median (indicating greater overlap in biomarker concentration’s distributions), the GAFAD model (open circle) achieved a higher AUC than GALAD (solid circle), highlighting its robustness in more challenging diagnostic settings. GAFAD, gender, age, alpha-fetoprotein (AFP), fucosylated alpha-fetoprotein (AFP-Fuc), and des-γ-carboxy prothrombin (DCP); GALAD, gender, age, AFP, Lens culinaris agglutinin-reactive alpha-fetoprotein (AFP-L3), and DCP; HCC, hepatocellular carcinoma.
Graphical abstract
GAFAD: A liquid chromatography-tandem mass spectrometry-based model for early hepatocellular carcinoma detection beyond GALAD’s limitations

Diagnostic performance of individual biomarkers and composite models for hepatocellular carcinoma (HCC) in the training and internal test sets

(1) HCC vs. non-HCC (2) Early-stage HCC (Stage II) vs. liver disease (3) Very-early-stage HCC (single node, ≤2 cm) vs. liver disease (4) AFP-negative HCC vs. liver disease
Cutoff value AUC (95% CI) Sensitivity (%) Specificity (%) Cutoff value AUC (95% CI) Sensitivity (%) Specificity (%) Cutoff value AUC (95% CI) Sensitivity (%) Specificity (%) Cutoff value AUC (95% CI) Sensitivity (%) Specificity (%)
AFP (established cutoff) 20 0.716 (0.669–0.760) 46.89 88.53 20 0.725 (0.641–0.809) 50.8 86.3 20 0.659 (0.596–0.718) 36.07 86.26 20 0.549 (0.486–0.609)
 AFP-Fuc% (Youden index) 8.3 0.736 (0.690–0.779) 48.59 94.95 8.3 0.733 (0.646–0.819) 49.2 94 8.3 0.616 (0.552–0.677) 32.79 94.51 8.3 0.634 (0.574–0.691) 27.66 94.51
 DCP (established cutoff) 40 0.846 (0.807–0.880) 61.02 98.17 40 0.855 (0.786–0.924) 67.8 97.3 40 0.716 (0.655–0.772) 29.51 97.8 40 0.826 (0.776–0.869) 53.19 97.8
 GALAD (established cutoff) −0.63 0.879 (0.843–0.910) 85.31 66.51 −0.63 0.877 (0.827–0.927) 89.8 59.9 −0.63 0.779 (0.722–0.830) 73.77 60.44 −0.63 0.741 (0.685–0.792) 72.34 60.44
 GALAD (Youden index) −0.03 0.879 (0.843–0.910) 79.66 76.61 0.48 0.877 (0.827–0.927) 76.3 81.3 1.27 0.779 (0.722–0.830) 49.18 92.86 −0.49 0.741 (0.685–0.792) 71.28 63.74
 GAFAD (Youden index) 0.11 0.937 (0.908–0.959) 78.53 95.41 0.64 0.933 (0.891–0.975) 78 97.8 0.11 0.880 (0.832–0.918) 63.93 94.51 0.11 0.882 (0.838–0.918) 64.89 94.51
 GAFA (Youden index) −0.03 0.866 (0.829–0.898) 71.19 83.03 0.26 0.851 (0.793–0.910) 71.2 85.7 −0.79 0.804 (0.748–0.852) 85.25 59.34 −1.2 0.742 (0.686–0.792) 91.49 47.8
Test set
 AFP (established cutoff) 20 0.747 (0.663–0.819) 46.55 87.5 20 0.713 (0.579–0.846) 42.9 82.8 20 0.698 (0.583–0.797) 42.11 84.48 20 0.513 (0.404–0.620)
 AFP-Fuc% (Youden index) 8.3 0.781 (0.700–0.848) 51.72 94.44 8.3 0.853 (0.752–0.954) 57.1 91.4 8.3 0.625 (0.508–0.733) 31.58 93.1 8.3 0.645 (0.537–0.744) 22.58 93.1
 DCP (established cutoff) 40 0.845 (0.772–0.903) 48.28 98.61 40 0.857 (0.746–0.969) 57.1 96.6 40 0.802 (0.696–0.884) 15.79 98.28 40 0.755 (0.653–0.840) 29.03 98.28
 GALAD (established cutoff) −0.63 0.910 (0.847–0.953) 91.38 75 −0.63 0.904 (0.839–0.968) 100 67.2 −0.63 0.803 (0.696–0.885) 73.68 68.97 −0.63 0.801 (0.703–0.878) 83.87 68.97
 GALAD (Youden index) −0.03 0.910 (0.847–0.953) 79.31 83.33 −0.46 0.904 (0.839–0.968) 100 72.4 1.25 0.803 (0.696–0.885) 52.63 94.83 −0.47 0.801 (0.703–0.878) 83.87 72.41
 GAFAD (Youden index) −0.69 0.947 (0.893–0.978) 91.38 88.89 −0.26 0.973 (0.940–1.000) 100 91.4 −0.69 0.888 (0.796–0.948) 84.21 86.21 −0.69 0.894 (0.811–0.950) 83.87 86.21
 GAFA (Youden index) −0.56 0.905 (0.841–0.949) 86.21 79.17 −0.18 0.916 (0.840–0.993) 90.5 82.8 −0.75 0.830 (0.727–0.906) 78.95 70.69 −0.56 0.821 (0.725–0.894) 80.65 74.14

Diagnostic performance was evaluated across multiple clinically relevant conditions, including the total cohort (HCC vs. non-HCC), early-stage HCC (mUICC stage II), very earlystage HCC (mUICC stage I), and AFP-negative HCC (AFP <20 ng/mL). For each biomarker or model, the cutoff value, area under the receiver operating characteristic curve (AUC) with 95% confidence interval (CI), sensitivity, and specificity are shown. “Established cutoff” values represent conventional clinical thresholds, whereas “Youden index” cutoffs were derived by maximizing Youden’s J statistic. Bold values indicate the highest AUC within each condition. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; mUICC, modified Union for International Cancer Control.

Table 1 Diagnostic performance of individual biomarkers and composite models for hepatocellular carcinoma (HCC) in the training and internal test sets

Diagnostic performance was evaluated across multiple clinically relevant conditions, including the total cohort (HCC vs. non-HCC), early-stage HCC (mUICC stage II), very earlystage HCC (mUICC stage I), and AFP-negative HCC (AFP <20 ng/mL). For each biomarker or model, the cutoff value, area under the receiver operating characteristic curve (AUC) with 95% confidence interval (CI), sensitivity, and specificity are shown. “Established cutoff” values represent conventional clinical thresholds, whereas “Youden index” cutoffs were derived by maximizing Youden’s J statistic. Bold values indicate the highest AUC within each condition. AFP, alpha-fetoprotein; AFP-Fuc, fucosylated alpha-fetoprotein; AFP-L3, Lens culinaris agglutinin-reactive alpha-fetoprotein; DCP, des-γ-carboxy prothrombin; GAFA, gender, age, AFP, and AFP-Fuc; GAFAD, gender, age, AFP, AFP-Fuc, and DCP; GALAD, gender, age, AFP, AFP-L3, and DCP; mUICC, modified Union for International Cancer Control.