| 本文已被:浏览 562次 下载 282次 |
 码上扫一扫! |
| 基于集成学习的沿海低能见度天气分类预报方法 |
|
陈锦鹏1,2,3, 林辉2,3, 吴雪菲4, 黄奕丹2,3, 程晶晶2,3, 庄毅斌2,3
|
|
1. 福建省灾害天气重点实验室,福建 福州 350001;2. 数字科学与统计重点实验室,福建 漳州 363005;3. 漳州市气象局,福建 漳州 363005;4. 福建省大气探测技术保障中心,福建 福州 350001
|
|
| 摘要: |
| 在2020年3月—2021年7月福建漳州沿海地区融合实况资料与欧洲中心细网格模式预报产品的基础上,应用集成学习中的LightGBM(Light Gradient Boosting Machine)算法建立分类预报模型以预测低能见度天气。针对样本极端不均衡的问题,在建模与检验中分别采用Bagging(Bootstrap Aggregating)技术和AUC(Area Under Curve)评分进行解决。根据有无新特征构造和模型融合划分为四种方案进行试验,同时将逻辑回归建模方案作为对比。结果表明:(1)在所有特征中,2 m露点对判断低能见度天气发生发展最为重要,2 m与1 000 hPa温差的重要性次之;(2)所有建模方案均能改善模式原始预报,其中LightGBM模型总体效果优于逻辑回归模型,两者命中率相似,但前者空报率显著降低;(3)新特征构造与模型融合的技巧能够进一步改善预测性能,包含这两者的建模方案在测试集上表现更佳,其中新特征构造对模型的提升幅度更为突出。 |
| 关键词: 低能见度 分类预报 集成学习 LoRa AUC |
| DOI:10.16032/j.issn.1004-4965.2023.059 |
| 分类号: |
| 基金项目: |
|
| CLASSIFICATION FORECAST METHOD OF COSTAL LOW VISIBILITY WEATHER BASED ON ENSEMBLE LEARNING |
|
CHEN Jinpeng1,2,3, LIN Hui2,3, WU Xuefei4, HUANG Yidan2,3, CHENG Jingjing2,3, ZHUANG Yibin2,3
|
|
1. Fujian Key Laboratory of Severe Weather, Fuzhou 350001, China;2. Fujian Key Laboratory of Data Science and Statistics, Zhangzhou, Fujian 363005, China;3. Zhangzhou Meteorological Bureau, Zhangzhou, Fujian 363005, China;4. Fujian Atmospheric Detection Technology Support Center, Fuzhou 350001, China
|
| Abstract: |
| A classification forecast method based on Light Gradient Boosting Machine (LightGBM) was utilized in this study to predict low visibility weather, using the coastal fusion observations and EC-thin model products of Zhangzhou from March 2020 to July 2021. The experiment was divided into four groups, including the new feature construction and model fusion schemes. The Bootstrap Aggregating (Bagging) technology and Area Under Curve (AUC) score were used to diminish the negative effect of extreme imbalance of samples, and the benchmark experiment employed the Logistics Regression (LR) method. The results showed that: (1) The most significant feature for estimating the possibility of low visibility weather was the 2?m dew point, followed by the temperature difference between 2m and 1000 hPa. (2) All model schemes exhibited improvement in comparison to the original forecast from the numerical model to varying degrees. In terms of metrics, the LightGBM model performed better than the LR model, largely due to its lower false alarm rate. (3) The skills of reasonable feature construction and model fusion contributed to optimizing the prediction performance and achieving higher scores on the test set. The impact of reasonable feature construction was particularly noteworthy. |
| Key words: low visibility classification forecast LightGBM LoRa AUC |