Factors Affecting Landslide Susceptibility Mapping: Assessing the Influence of Different Machine Learning Approaches, Sampling Strategies and Data Splitting

Minu Treesa Abraham; Neelima Satyam; Revuri Lokesh; Biswajeet Pradhan; Abdullah Alamri

doi:10.3390/land10090989

Factors Affecting Landslide Susceptibility Mapping: Assessing the Influence of Different Machine Learning Approaches, Sampling Strategies and Data Splitting

Land ◽

10.3390/land10090989 ◽

2021 ◽

Vol 10 (9) ◽

pp. 989

Author(s):

Minu Treesa Abraham ◽

Neelima Satyam ◽

Revuri Lokesh ◽

Biswajeet Pradhan ◽

Abdullah Alamri

Keyword(s):

Machine Learning ◽

Landslide Susceptibility ◽

Sampling Strategy ◽

Susceptibility Mapping ◽

Landslide Susceptibility Mapping ◽

Support Vector ◽

Susceptibility Map ◽

Learning Approaches ◽

Sampling Strategies ◽

Data Splitting

Data driven methods are widely used for the development of Landslide Susceptibility Mapping (LSM). The results of these methods are sensitive to different factors, such as the quality of input data, choice of algorithm, sampling strategies, and data splitting ratios. In this study, five different Machine Learning (ML) algorithms are used for LSM for the Wayanad district in Kerala, India, using two different sampling strategies and nine different train to test ratios in cross validation. The results show that Random Forest (RF), K Nearest Neighbors (KNN), and Support Vector Machine (SVM) algorithms provide better results than Naïve Bayes (NB) and Logistic Regression (LR) for the study area. NB and LR algorithms are less sensitive to the sampling strategy and data splitting, while the performance of the other three algorithms is considerably influenced by the sampling strategy. From the results, both the choice of algorithm and sampling strategy are critical in obtaining the best suited landslide susceptibility map for a region. The accuracies of KNN, RF, and SVM algorithms have increased by 10.51%, 10.02%, and 4.98% with the use of polygon landslide inventory data, while for NB and LR algorithms, the performance was slightly reduced with the use of polygon data. Thus, the sampling strategy and data splitting ratio are less consequential with NB and algorithms, while more data points provide better results for KNN, RF, and SVM algorithms.

Download Full-text

A Novel Hybrid Method for Landslide Susceptibility Mapping-Based GeoDetector and Machine Learning Cluster: A Case of Xiaojin County, China

ISPRS International Journal of Geo-Information ◽

10.3390/ijgi10020093 ◽

2021 ◽

Vol 10 (2) ◽

pp. 93

Author(s):

Wei Xie ◽

Xiaoshuang Li ◽

Wenbin Jian ◽

Yang Yang ◽

Hongwei Liu ◽

...

Keyword(s):

Machine Learning ◽

Hybrid Method ◽

Roc Curve ◽

Landslide Susceptibility ◽

Susceptibility Mapping ◽

Assessment Model ◽

Landslide Susceptibility Mapping ◽

Support Vector ◽

Susceptibility Map ◽

Area Index

Landslide susceptibility mapping (LSM) could be an effective way to prevent landslide hazards and mitigate losses. The choice of conditional factors is crucial to the results of LSM, and the selection of models also plays an important role. In this study, a hybrid method including GeoDetector and machine learning cluster was developed to provide a new perspective on how to address these two issues. We defined redundant factors by quantitatively analyzing the single impact and interactive impact of the factors, which was analyzed by GeoDetector, the effect of this step was examined using mean absolute error (MAE). The machine learning cluster contains four models (artificial neural network (ANN), Bayesian network (BN), logistic regression (LR), and support vector machines (SVM)) and automatically selects the best one for generating LSM. The receiver operating characteristic (ROC) curve, prediction accuracy, and the seed cell area index (SCAI) methods were used to evaluate these methods. The results show that the SVM model had the best performance in the machine learning cluster with the area under the ROC curve of 0.928 and with an accuracy of 83.86%. Therefore, SVM was chosen as the assessment model to map the landslide susceptibility of the study area. The landslide susceptibility map demonstrated fit with landslide inventory, indicated the hybrid method is effective in screening landslide influences and assessing landslide susceptibility.

Download Full-text

Landslide susceptibility mapping based on convolutional neural network and conventional machine learning methods

10.21203/rs.3.rs-190195/v1 ◽

2021 ◽

Author(s):

Rui Liu ◽

Xin Yang ◽

Chong Xu ◽

Luyao Li ◽

Xiangqiang Zeng

Keyword(s):

Neural Network ◽

Machine Learning ◽

Convolutional Neural Network ◽

Landslide Susceptibility ◽

Susceptibility Mapping ◽

Landslide Susceptibility Mapping ◽

Support Vector ◽

Learning Methods ◽

Machine Learning Methods ◽

Conventional Machine

Abstract Landslide susceptibility mapping (LSM) is a useful tool to estimate the probability of landslide occurrence, providing a scientific basis for natural hazards prevention, land use planning, and economic development in landslide-prone areas. To date, a large number of machine learning methods have been applied to LSM, and recently the advanced Convolutional Neural Network (CNN) has been gradually adopted to enhance the prediction accuracy of LSM. The objective of this study is to introduce a CNN based model in LSM and systematically compare its overall performance with the conventional machine learning models of random forest, logistic regression, and support vector machine. Herein, we selected the Jiuzhaigou region in Sichuan Province, China as the study area. A total number of 710 landslides and 12 predisposing factors were stacked to form spatial datasets for LSM. The ROC analysis and several statistical metrics, such as accuracy, root mean square error (RMSE), Kappa coefficient, sensitivity, and specificity were used to evaluate the performance of the models in the training and validation datasets. Finally, the trained models were calculated and the landslide susceptibility zones were mapped. Results suggest that both CNN and conventional machine-learning based models have a satisfactory performance (AUC: 85.72% − 90.17%). The CNN based model exhibits excellent good-of-fit and prediction capability, and achieves the highest performance (AUC: 90.17%) but also significantly reduces the salt-of-pepper effect, which indicates its great potential of application to LSM.

Download Full-text

Landslide susceptibility mapping using machine learning for Wenchuan County, Sichuan province, China

E3S Web of Conferences ◽

10.1051/e3sconf/202019803023 ◽

2020 ◽

Vol 198 ◽

pp. 03023

Author(s):

Xin Yang ◽

Rui Liu ◽

Luyao Li ◽

Mei Yang ◽

Yuantao Yang

Keyword(s):

Machine Learning ◽

Landslide Susceptibility ◽

Susceptibility Mapping ◽

Machine Learning Algorithms ◽

Landslide Susceptibility Mapping ◽

Support Vector ◽

Roc Curve Analysis ◽

Learning Methods ◽

Machine Learning Methods ◽

Boosted Decision Tree

Landslide susceptibility mapping is a method used to assess the probability and spatial distribution of landslide occurrences. Machine learning methods have been widely used in landslide susceptibility in recent years. In this paper, six popular machine learning algorithms namely logistic regression, multi-layer perceptron, random forests, support vector machine, Adaboost, and gradient boosted decision tree were leveraged to construct landslide susceptibility models with a total of 1365 landslide points and 14 predisposing factors. Subsequently, the landslide susceptibility maps (LSM) were generated by the trained models. LSM shows the main landslide zone is concentrated in the southeastern area of Wenchuan County. The result of ROC curve analysis shows that all models fitted the training datasets and achieved satisfactory results on validation datasets. The results of this paper reveal that machine learning methods are feasible to build robust landslide susceptibility models.

Download Full-text

Comparison between Deep Learning and Tree-Based Machine Learning Approaches for Landslide Susceptibility Mapping

Water ◽

10.3390/w13192664 ◽

2021 ◽

Vol 13 (19) ◽

pp. 2664

Author(s):

Sunil Saha ◽

Jagabandhu Roy ◽

Tusar Kanti Hembram ◽

Biswajeet Pradhan ◽

Abhirup Dikshit ◽

...

Keyword(s):

Neural Network ◽

Machine Learning ◽

Deep Learning ◽

Landslide Susceptibility ◽

Learning Model ◽

Susceptibility Mapping ◽

Landslide Susceptibility Mapping ◽

Learning Approaches ◽

Statistical Measures ◽

Deep Learning Model

The efficiency of deep learning and tree-based machine learning approaches has gained immense popularity in various fields. One deep learning model viz. convolution neural network (CNN), artificial neural network (ANN) and four tree-based machine learning models, namely, alternative decision tree (ADTree), classification and regression tree (CART), functional tree and logistic model tree (LMT), were used for landslide susceptibility mapping in the East Sikkim Himalaya region of India, and the results were compared. Landslide areas were delimited and mapped as landslide inventory (LIM) after gathering information from historical records and periodic field investigations. In LIM, 91 landslides were plotted and classified into training (64 landslides) and testing (27 landslides) subsets randomly to train and validate the models. A total of 21 landslide conditioning factors (LCFs) were considered as model inputs, and the results of each model were categorised under five susceptibility classes. The receiver operating characteristics curve and 21 statistical measures were used to evaluate and prioritise the models. The CNN deep learning model achieved the priority rank 1 with area under the curve of 0.918 and 0.933 by using the training and testing data, quantifying 23.02% and 14.40% area as very high and highly susceptible followed by ANN, ADtree, CART, FTree and LMT models. This research might be useful in landslide studies, especially in locations with comparable geophysical and climatological characteristics, to aid in decision making for land use planning.

Download Full-text

Landslide Susceptibility Mapping Using the Stacking Ensemble Machine Learning Method in Lushui, Southwest China

Applied Sciences ◽

10.3390/app10114016 ◽

2020 ◽

Vol 10 (11) ◽

pp. 4016 ◽

Cited By ~ 3

Author(s):

Xudong Hu ◽

Han Zhang ◽

Hongbo Mei ◽

Dunhui Xiao ◽

Yuanyuan Li ◽

...

Keyword(s):

Machine Learning ◽

Landslide Susceptibility ◽

Southwest China ◽

Susceptibility Mapping ◽

Landslide Susceptibility Mapping ◽

Support Vector ◽

Machine Learning Method ◽

Learning Method ◽

Statistical Measures ◽

Ensemble Machine Learning

Landslide susceptibility mapping is considered to be a prerequisite for landslide prevention and mitigation. However, delineating the spatial occurrence pattern of the landslide remains a challenge. This study investigates the potential application of the stacking ensemble learning technique for landslide susceptibility assessment. In particular, support vector machine (SVM), artificial neural network (ANN), logical regression (LR), and naive Bayes (NB) were selected as base learners for the stacking ensemble method. The resampling scheme and Pearson’s correlation analysis were jointly used to evaluate the importance level of these base learners. A total of 388 landslides and 12 conditioning factors in the Lushui area (Southwest China) were used as the dataset to develop landslide modeling. The landslides were randomly separated into two parts, with 70% used for model training and 30% used for model validation. The models’ performance was evaluated using the area under the receiver operating characteristic (ROC) curve (AUC) and statistical measures. The results showed that the stacking-based ensemble model achieved an improved predictive accuracy as compared to the single algorithms, while the SVM-ANN-NB-LR (SANL) model, the SVM-ANN-NB (SAN) model, and the ANN-NB-LR (ANL) models performed equally well, with AUC values of 0.931, 0.940, and 0.932, respectively, for validation stage. The correlation coefficient between the LR and SVM was the highest for all resampling rounds, with a value of 0.72 on average. This connotes that LR and SVM played an almost equal role when the ensemble of SANL was applied for landslide susceptibility analysis. Therefore, it is feasible to use the SAN model or the ANL model for the study area. The finding from this study suggests that the stacking ensemble machine learning method is promising for landslide susceptibility mapping in the Lushui area and is capable of targeting areas prone to landslides.

Download Full-text

Hybrid ensemble machine learning approaches for landslide susceptibility mapping using different sampling ratios at East Sikkim Himalayan, India

Advances in Space Research ◽

10.1016/j.asr.2021.05.018 ◽

2021 ◽

Author(s):

Sunil Saha ◽

Jagabandhu Roy ◽

Biswajeet Pradhan ◽

Tusar Kanti Hembram

Keyword(s):

Machine Learning ◽

Landslide Susceptibility ◽

Susceptibility Mapping ◽

Landslide Susceptibility Mapping ◽

Learning Approaches ◽

Ensemble Machine Learning

Download Full-text

Landslide susceptibility mapping using Forest by Penalizing Attributes (FPA) algorithm based machine learning approach

Vietnam Journal of Earth Sciences ◽

10.15625/0866-7187/42/3/15047 ◽

2020 ◽

Vol 42 (3) ◽

Cited By ~ 1

Author(s):

Tran Van Phong ◽

Hai-Bang Ly ◽

Phan Trong Trinh ◽

Indra Prakash

Keyword(s):

Machine Learning ◽

Landslide Susceptibility ◽

Susceptibility Mapping ◽

Landslide Susceptibility Mapping ◽

Susceptibility Map ◽

Conditioning Factors ◽

Environmental Conditioning ◽

Landslide Modeling ◽

Machine Learning Approach ◽

First Time

Landslide susceptibility mapping is a helpful tool for assessment and management of landslides of an area. In this study, we have applied first time Forest by Penalizing Attributes (FPA) algorithm-based Machine Learning (ML) approach for mapping of landslide susceptibility at Muong Lay district (Vietnam). For this aim, 217 historical landslides locations were identified and analyzed for the development of FPA model and generation of susceptibility map. Nine landslide topographical and geo-environmental conditioning factors (curvature, geology/lithology, aspect, distance from faults, rivers and roads, weathering crust, slope, and deep division) were utilized to construct the training and validating datasets for landslide modeling. Different quantitative statistical indices including Area Under the Receiver Operating Characteristic (ROC) curve (AUC) were used to evaluate the performance of the model. The results indicate that the predictive capability of the FPA is very good for landslide susceptibility mapping on both training (AUC = 0.935) and validating (AUC = 0.882) datasets. Thus, the novel FPA based ML model can be utilized for the development of accurate landslide susceptibility map of the study area and this approach can also be applied in other landslide prone areas.

Download Full-text

Conditioning Factors Determination for Landslide Susceptibility Mapping Using Support Vector Machine Learning

IGARSS 2019 - 2019 IEEE International Geoscience and Remote Sensing Symposium ◽

10.1109/igarss.2019.8898340 ◽

2019 ◽

Author(s):

Bahareh Kalantar ◽

Naonori Ueda ◽

Usman Salihu Lay ◽

Husam Abdulrasool H. Al-Najjar ◽

Alfian Abdul Halin

Keyword(s):

Machine Learning ◽

Support Vector Machine ◽

Landslide Susceptibility ◽

Susceptibility Mapping ◽

Landslide Susceptibility Mapping ◽

Support Vector ◽

Conditioning Factors

Download Full-text

Decision tree based ensemble machine learning approaches for landslide susceptibility mapping

Geocarto International ◽

10.1080/10106049.2021.1892210 ◽

2021 ◽

pp. 1-35

Author(s):

Alireza Arabameri ◽

Subodh Chandra Pal ◽

Fatemeh Rezaie ◽

Rabin Chakrabortty ◽

Asish Saha ◽

...

Keyword(s):

Machine Learning ◽

Decision Tree ◽

Landslide Susceptibility ◽

Susceptibility Mapping ◽

Landslide Susceptibility Mapping ◽

Learning Approaches ◽

Ensemble Machine Learning

Download Full-text

Using landslide-inventory mapping for a combined bagged-trees and logistic-regression approach to determining landslide susceptibility in eastern Kentucky, United States

Quarterly Journal of Engineering Geology and Hydrogeology ◽

10.1144/qjegh2020-177 ◽

2021 ◽

pp. qjegh2020-177

Author(s):

Matthew M. Crawford ◽

Jason M. Dortch ◽

Hudson J. Koch ◽

Ashton A. Killen ◽

Junfeng Zhu ◽

...

Keyword(s):

Machine Learning ◽

Logistic Regression ◽

Standard Deviation ◽

Landslide Susceptibility ◽

Engineering Geology ◽

Susceptibility Mapping ◽

Landslide Inventory ◽

Landslide Susceptibility Mapping ◽

Susceptibility Map ◽

Landslide Occurrence

High-resolution LiDAR-derived datasets from a 1.5-m digital elevation model and a detailed landslide inventory (N ≥ 1,000) for Magoffin County, Kentucky, USA, were used to develop a combined machine-learning and statistical approach to improve geomorphic-based landslide-susceptibility mapping.An initial dataset of 36 variables was compiled to investigate the connection between slope morphology and landslide occurrence. Bagged trees, a machine-learning random-forest classifier, was used to evaluate the geomorphic variables, and 12 were identified as important: standard deviation of plan curvature, standard deviation of elevation, sum of plan curvature, minimum slope, mean plan curvature, range of elevation, sum of roughness, mean curvature, sum of curvature, mean roughness, minimum curvature, and standard deviation of curvature. These variables were further evaluated using logistic regression to determine the probability of landslide occurrence and then used to create a landslide-susceptibility map.The performance of the logistic-regression model was evaluated by the receiver operating characteristic curve, area under the curve, which was 0.83. Standard deviations from the probability mean were used to set landslide-susceptibility classifications: low (0–0.10), low–moderate (0.11–0.27), moderate (0.28–0.44), moderate–high (0.45–0.7), and high (0.7–1.0). Logistic-regression results were validated by using a separate landslide inventory for the neighboring Prestonsburg 7.5-minute quadrangle, and running the same regression function. Results indicate that 74.9 percent of the landslide deposits were identified as having moderate, moderate–high, or high landslide susceptibility. Combining inventory mapping with statistical modelling identified important geomorphic variables and produced a useful approach to landslide-susceptibility mapping.Thematic collection: This article is part of the Digitization and Digitalization in engineering geology and hydrogeology collection available at: https://www.lyellcollection.org/cc/digitization-and-digitalization-in-engineering-geology-and-hydrogeology

Download Full-text