Abstract:
Objective To construct a risk prediction model for prolonged hospital stay (LOS) in patients with liver cancer after hepatectomy based on interpretable machine learning algorithms.
Methods The clinical data of 527 patients with liver cancer who underwent liver resection were retrospectively collected, and randomly divided into the training set (370 cases) and validation set (157 cases) at a ratio of 7:3. The common feature variables were screened through Lasso regression and Boruta algorithm, and six machine learning algorithms, namely logistic regression (LR), Support Vector machine (SVM), Gradient boosting machine (GBM), Extreme gradient boosting (XGBoost), K-Nearest Neighbor (KNN) and Adaptive boosting (Adaboost), were respectively used to establish the prediction models, and proceed simultaneously cross verification at a 10% discount. The area under the receiver operating characteristic curve (AUC) was adopted to select the optimal model, and the Shapley additive interpretation (SHAP) method was used to analyze the contribution of characteristic variables.
Result Among the 527 patients, the incidence of prolonged hospital stay was 23.91% (126/527). After feature screening, seven feature variables were determined for constructing the machine learning model. Among them, the KNN model had the best predictive performance, with an AUC of 0.946 (95%CI: 0.925–0.967). The SHAP bar chart showed that the order of importance was the resection range, albumin-bilirubin grade, intraoperative blood loss, intraoperative infusion volume, maximum tumor diameter, albumin, and operation time in turn.
Conclusions The KNN model has the best efficacy in predicting the prolonged hospital stay of liver cancer patients after hepatectomy, which is conducive to the early identification of high-risk groups by medical staff and implementation of targeted management measures.