| Korean J Health Promot > Volume 26(3); 2026 > Article |
|
AUTHOR CONTRIBUTIONS
Dr. Jeong Yun PARK had full access to all of the data in the study and takes responsibility for the integrity of the data and the accuracy of the data analysis. All authors reviewed this manuscript and agreed to individual contributions.
Conceptualization: SML and JYP. Data curation: SML. Formal analysis: SML and JYP. Investigation: SML. Methodology: SML and JYP. Software: SML. Validation: SML and JYP. Writing–original draft: SML. Writing–review & editing: JYP.
| Study (ref.) | Population | Development/validation | Events/total (unit, %) | Outcome | Candidate/final predictors | Selection method | Final model variables | Model type | Discrimination (AUC/C-index, 95% CI) | Validation; interpretability |
|---|---|---|---|---|---|---|---|---|---|---|
| Models targeting removal itself (direct evidence, n=3) | ||||||||||
| Zhang et al. (2023) [17] | Cancer patients, China | Development 3,391 (train 2,374/test 1,017); external validation 600 | 284/3,391 (per patient, 8.38%) | Unplanned PICC removal: early removal for a severe complication, or accidental dislodgement | 33 → 10 | Univariable screening, LASSO (11 variables), then multivariable backward logistic regression | Impaired physical mobility; elevated D-dimer; diabetes; surgical history; more than one puncture; surgical treatment; targeted therapy. Protective: valved catheter, normal BMI, polyurethane material | SVM (compared with logistic regression and random forest; TOPSIS used for final selection) | 0.904 (train)/0.875 (test)/0.718 (external validation) | External (two hospitals) and internal; Hosmer–Lemeshow P=0.06; DCA reported. |
| Yang et al. (2026) [18] | Cancer patients, outpatient catheter-maintenance clinic, China | 212 (23 events/189 non-events) | 23/212 (per patient, 10.8%) | Unplanned PICC removal requiring new vascular access | 30 → 13 significant on univariable screening (P<0.100) | Univariable screening; weight-of-evidence encoding; focal loss for class imbalance | A composite index of health-related quality of life and self-management ability ranked first by SHAP, followed by self-management score and mid-upper-arm circumference of the PICC arm | XGBoost with focal loss (compared with random forest, SVM and logistic regression) | AUC 0.994; recall 0.870; accuracy 0.967; F1 0.851 | Internal (stratified 5-fold CV); Brier score 0.03, calibration plot and DCA; SHAP. External validation not performed. |
| Lee and Park (2025) [22] | Patients receiving parenteral nutrition, Korea | 218 procedures (train 174/test 44) | 18/218 (per procedure, 8.3%) | Complication-induced PICC removal (infection, occlusion, exit-site discomfort or dislodgement), analysed as time to event | 10 → 10 | Prespecified candidate set; no variable selection | Sex; age; cancer diagnosis; previous central venous catheter; first procedure; insertion side; insertion vein; insertion length; catheter diameter; interval from admission to request | Logistic regression, SVM, random forest and XGBoost for classification; DeepSurv, DeepHit and random survival forest for time to event | C-index: DeepSurv 0.611 (0.158–0.896); DeepHit 0.475; random survival forest 0.329. Mean classification accuracy 0.92 against an uninformative baseline of 91.7% | Internal (5-fold CV; separate test set of 44); integrated Brier score reported. External validation not performed. |
| Models targeting removal-precipitating complications (indirect evidence, n=11) | ||||||||||
| Sheng and Gao (2024) [19] | Adults with PICC, single hospital catheter clinic, China | 1,065 patients (80:20 split) | 76/1,065 (per patient, 7.14%) | PICC-related deep vein thrombosis confirmed on Doppler ultrasound after weekly Constans score and D-dimer screening | 21 → 21 | None; all 21 literature-derived variables were entered | Catheter-to-vein ratio ranked first in all three algorithms (random forest importance 91.55%) | Random forest, artificial neural network, support vector classifier | RF 0.86/ANN 0.81/SVC 0.77 (mean 0.81) | Internal (10-fold CV). The reported precision, recall and accuracy take “no DVT”, the majority class in 92.86% of the sample, as the positive label. |
| Li et al. (2024) [23] | Adults with PICC, single municipal hospital, China | 505 patients | 75/505 (per patient, 14.85%) | PICC-related infection | NR → 7 | LASSO followed by multivariable logistic regression | Age >60 yr; catheter movement; maintenance interval >7 days; direct puncture; impaired immunity; concurrent complication; pre-insertion temperature ≥37.2 °C | Logistic regression/nomogram | 0.889 | Internal. Several odds ratios lie close to the lower confidence limit, so internal reporting consistency is uncertain. |
| Song et al. (2020) [24] | Cancer patients on a first PICC for chemotherapy, China | 339 patients | 59/339 (per patient, 17.4%) | PICC-associated thrombosis on routine ultrasound 2 weeks after insertion; 96.6% asymptomatic and 93.2% fibrin sleeve | NR → 4 | Fisher scoring followed by multivariable logistic regression | Performance status; number of lumens; D-dimer; height (protective). Catheter-to-vein ratio was significant on univariable testing only and was explicitly excluded by the authors. | Logistic regression/nomogram | 0.822±0.031 for the best five-variable combination; 0.780±0.033 for the two-variable model the authors designated as best | No validation (apparent performance only; no split, cross-validation or bootstrap). |
| Chopra et al. (2017) [25] | Hospitalised adults on general medicine wards or in ICU, 51-hospital consortium, USA | 23,010 patients | 475/23,010 (per patient, 2.1%) | Symptomatic, imaging-confirmed upper-extremity deep vein thrombosis; imaging was performed only when symptoms were present. | NR → 5 | Univariable mixed-effects screening (P≤0.10) then stepwise selection by the Schwarz criterion; 10-fold multiple imputation | Another central venous catheter (1 point); WBC >12,000 (1 point); active cancer (2 points); two lumens (2 points) or three to four lumens (3 points); venous thromboembolism history (2 points) or within 30 days (3 points) | Point-based risk score derived from a logistic mixed model | 0.71 for the mixed model (estimated optimism 0.04); 0.65 on marginal predicted probability (optimism 0.03) | Internal only (200 bootstrap replications; calibration intercept 0.35, slope 0.90 [0.78–1.14]). External validation was not performed in this study. |
| Xia et al. (2026) [26] | Critically ill adults, MIMIC-IV database, USA | 8,145 insertions/6,376 patients (train 6,516/test 1,629) | 1,375/8,145 (per insertion, 16.88%) | PICC-associated thrombotic complications identified from ICD-9 and ICD-10 codes during the same hospitalisation | 29 → 29 | Prespecified candidate set; no selection. SMOTE was applied to the training set only, after the split. | All 29 variables retained. Prior thrombosis history was dominant (Gini importance 0.138), followed by INR (0.067). No catheter variables were available in the database. | Random forest (compared with logistic regression, SVM, gradient boosting and XGBoost) | 0.809 (0.780–0.838); 0.696 under patient-clustered splitting; 0.717 restricted to each patient’s first PICC | Internal (5-fold CV; bootstrap CI; DeLong test). TRIPOD+AI checklist reported; Gini importance and SHAP. External validation not performed. |
| Su et al. (2026) [27] | Patients with haematological malignancies, multicentre, China | 4,015 (train 2,810/independent test 1,205) | 232/4,015 (per patient, 5.8%) | Symptomatic PICC-related thrombosis | NR → 9, of which 5 were independent predictors | LASSO | History of venous thromboembolism; triple-lumen catheter; immunomodulatory drug use; peak D-dimer (risk rising sharply above about 1.5 mg/L); peak neutrophil-to-lymphocyte ratio | XGBoost (compared with a logistic regression baseline) | 0.862 (0.825–0.899) in the independent test set | Independent test set; SHAP; open web calculator. Negative predictive value 98.6% and positive predictive value 18.8% |
| Cui et al. (2026) [28] | Patients with gastrointestinal malignancies undergoing interventional therapy, China | 782 patients | 74/782 (per patient, cumulative incidence 7.8% at 1 month and 9.2% at 3 months) | Symptomatic PICC-related venous thrombosis, with death and non-thrombotic unplanned removal treated as competing events | 7 → 4 | Fine–Gray sub-distribution regression with AIC-based backward elimination | Catheter-to-vein ratio>0.45; non-cavoatrial tip position; elevated baseline D-dimer; neutrophil-to-lymphocyte ratio≥3.0 | Competing-risk nomogram | C-index 0.75 | Internal (1,000 bootstrap replications); calibration curves at 1, 2 and 3 months |
| Guo et al. (2025) [29] | Patients with haematological malignancies, single tertiary centre, China | 764 (train 534/validation 230) | 46/764 (per patient, 6.02%) | PICC catheter-related bloodstream infection (concordant peripheral and catheter-tip cultures, or differential time to positivity) | 18 → 9 | LASSO with 10-fold CV then stepwise multivariable logistic regression, both fitted on SMOTE-NC-augmented training data | Diabetes history; age; dwell time>60 days; two or more insertion attempts; dual-lumen catheter; cephalic vein; absolute neutrophil count<1.5×10⁹/L; absolute lymphocyte count<1.0×10⁹/L; D-dimer≥0.5 mg/L | Logistic regression/nomogram | 0.883 (0.863–0.903) in training and 0.822 (0.719–0.924) in validation | Internal validation set; calibration with bootstrap resampling; DCA. Estimates and training discrimination derive from oversampled data |
| Li et al. (2025) [31] | Adults with PICC, 27 hospitals, China | 3,453 patients (train 2,764/test 691) | 525/3,453 (per patient, 15.2%), of which 401 (76.4%) led to catheter removal | PICC-related venous thrombosis as time to event (symptoms combined with ultrasound findings) | About 48 → 26 (a 16-variable model was also built) | Univariable logistic screening carried out on the whole data set before the train–test split | Two prespecified sets of 26 and 16 variables | DeepSurv, DeepHit and Cox-Time; MP-RSF, MP-AdaBoost, ThresReg and MP-LogitR | 26-variable C-index: DeepSurv 0.95, Cox-Time 0.949, DeepHit 0.948, versus 0.707–0.772 for the conventional algorithms. The same DeepSurv model falls to 0.759 with 16 variables | Internal (5-fold CV); integrated Brier score; intraclass correlation for stability |
| Gao et al. (2025) [32] | Adult cancer inpatients, single cancer hospital, China | 20,941 patients (train 14,659/validation 6,282) | 227/20,941 (per patient, 1.1%) | Catheter occlusion, defined as inability to flush with saline or to aspirate blood | More than 90 → 59 → 13 | LASSO (λ=0.001067, 10-fold CV) | Catheter type; insertion length; dwell days; sex; electrolyte disturbance; primary tumour site; thrombophilia; cough; malnutrition; number of insertion attempts; number of chemotherapy agents; herbal medicine; interleukin | XGBoost (compared with logistic regression and random forest) | Training: RF 0.976/XGBoost 0.929/logistic 0.786. Validation: logistic 0.773/XGBoost 0.759/RF 0.643 | Internal; SHAP. Logistic regression outperformed both tree ensembles in validation, yet XGBoost was declared the optimal model. |
| Hu et al. (2025) [30] | Cancer patients, prospective cohort, China | 281 enrolled/275 analysed | 18/275 (per patient, 6.5%) | Symptomatic PICC-related venous thrombosis | NR → 4 | Univariable screening then stepwise multivariable logistic regression | Insulin-requiring diabetes 8.016 (1.157–55.536); major surgery; reduced limb activity of the PICC arm; catheter material | Logistic regression/nomogram | 0.796 (0.695–0.897) | Internal; Hosmer–Lemeshow 1.685, P=0.194. External validation not performed. |
Not-reported items are entered as NR. In [26] the primary estimate of 0.809 falls to 0.696 when the data are split by patient rather than by insertion, indicating that part of the apparent discrimination reflects repeated insertions in the same patient. In [29] both variable selection and model fitting were carried out on SMOTE-NC-augmented training data, so the reported precision is optimistic relative to the 46 observed events. In [31] variable selection preceded the train–test split. In [24] no form of internal validation was reported. In [19] the reported precision, recall and accuracy take the majority class as the positive label. In [22] the mean classification accuracy of 0.92 is indistinguishable from the uninformative baseline of 91.7% implied by an event rate of 8.3%.
AIC, Akaike information criterion; ANN, artificial neural network; AUC, area under the curve; BMI, body mass index; CI, confidence interval; CV, cross-validation; DCA, decision curve analysis; DVT, deep vein thrombosis; ICD, International Classification of Diseases; ICU, intensive care unit; INR, international normalised ratio; LASSO, least absolute shrinkage and selection operator; MIMIC, Medical Information Mart for Intensive Care; NR, not reported; PICC, peripherally inserted central catheter; ref., reference; RF, random forest; SHAP, SHapley Additive exPlanations; SMOTE, synthetic minority over-sampling technique; SMOTE-NC, SMOTE for nominal and continuous features; SVC, support vector classifier; SVM, support vector machine; TOPSIS, technique for order of preference by similarity to ideal solution; TRIPOD, transparent reporting of a multivariable prediction model for individual prognosis or diagnosis; WBC, white blood cell.
Seon Min LEE
https://orcid.org/0009-0004-8694-2169
Jeong Yun PARK
https://orcid.org/0000-0002-0210-8213