基于可解释性机器学习构建肝癌病人肝切除术后住院时间延长风险预测模型

    Construction on a risk prediction model for prolonged hospital stay in patients with liver cancer after hepatectomy based on interpretable machine learning

    • 摘要:
      目的: 基于可解释性机器学习算法构建肝癌病人肝切除术后住院时间(LOS)延长风险预测模型。
      方法: 回顾性收集527例肝癌肝切除病人的临床资料,按7∶3的比例随机划分为训练集370例和验证集157例。通过Lasso回归和Boruta算法筛选共同特征变量,并运用逻辑回归(LR)、支持向量机(SVM)、梯度提升机(GBM)、极限梯度提升(XGBoost)、K最近邻(KNN)及自适应提升(Adaboost)6种机器学习算法分别建立预测模型并进行10折交叉验证。采用受试者工作特征曲线下面积(AUC)选择最优模型,使用Shapley加性解释(SHAP)法解析特征变量的贡献度。
      结果: 527例病人中,住院时间延长发生率为23.91%(126/527)。经特征筛选后确定7个特征变量用于构建机器学习模型,其中KNN模型预测性能最佳,AUC为0.946(95%CI:0.925 ~ 0.967)。SHAP条形图显示重要性排序为切除范围、白蛋白–胆红素分级、术中失血量、术中输液量、最大肿瘤直径、白蛋白和手术时间。
      结论: 基于KNN模型预测肝癌病人肝切除术后住院时间延长的效能最优,有利于医护人员早期识别高危人群并给予针对性管理措施。

       

      Abstract:
      Objective To construct a risk prediction model for prolonged hospital stay (LOS) in patients with liver cancer after hepatectomy based on interpretable machine learning algorithms.
      Methods The clinical data of 527 patients with liver cancer who underwent liver resection were retrospectively collected, and randomly divided into the training set (370 cases) and validation set (157 cases) at a ratio of 7:3. The common feature variables were screened through Lasso regression and Boruta algorithm, and six machine learning algorithms, namely logistic regression (LR), Support Vector machine (SVM), Gradient boosting machine (GBM), Extreme gradient boosting (XGBoost), K-Nearest Neighbor (KNN) and Adaptive boosting (Adaboost), were respectively used to establish the prediction models, and proceed simultaneously cross verification at a 10% discount. The area under the receiver operating characteristic curve (AUC) was adopted to select the optimal model, and the Shapley additive interpretation (SHAP) method was used to analyze the contribution of characteristic variables.
      Result Among the 527 patients, the incidence of prolonged hospital stay was 23.91% (126/527). After feature screening, seven feature variables were determined for constructing the machine learning model. Among them, the KNN model had the best predictive performance, with an AUC of 0.946 (95%CI: 0.925–0.967). The SHAP bar chart showed that the order of importance was the resection range, albumin-bilirubin grade, intraoperative blood loss, intraoperative infusion volume, maximum tumor diameter, albumin, and operation time in turn.
      Conclusions The KNN model has the best efficacy in predicting the prolonged hospital stay of liver cancer patients after hepatectomy, which is conducive to the early identification of high-risk groups by medical staff and implementation of targeted management measures.

       

    /

    返回文章
    返回