| 摘要: |
| 【目的】 构建一套兼顾统计规律与专业机理的多维协同数据清洗框架,提升耕地质量数据的可靠性与可用性。【方法】 文章以江西省“国—省—市—县”4级监测网的长时序数据为基础,在预处理阶段引入基于领域知识的业务逻辑约束进行初步清洗,继而构建涵盖统计分布(3σ原则、IQR法)、时序特征(卡尔曼滤波、移动平均)和空间结构(半变异函数、局部莫兰指数)的三维并行检测矩阵,并引入集成投票机制对多维度识别结果进行交叉验证与综合决策。【结果】 (1)该框架在土壤属性与作物产量数据中表现出良好的鲁棒性,能有效识别统计、时空及逻辑等多类型异常;(2)集成投票机制在异常识别与信息保全之间实现了动态平衡,既剔除了高置信度异常又保留了真实高值,克服了传统单一方法的局限性;(3)领域知识在预处理阶段的引入显著提升了清洗深度,通过逻辑校验机制识别出统计方法无法发现的复合型错误。【结论】 该研究提出的多维度协同清洗框架在实现识别异常与保全信息动态平衡方面具有显著优势,为农业大数据的源头治理与高精度分析提供了一套可复用的技术方案。 |
| 关键词: 耕地质量 数据清洗 时空异质性 集成投票机制 |
| DOI:10.12105/j.issn.1672-0423.20250410 |
| 分类号: |
| 基金项目:国家重点研发项目“坡耕地红壤控酸决策支持系统开发及应用”(2022YFD1900601-4) |
|
| Multi-source heterogeneous cultivated land quality data anomaly identification and cleaning strategy based on ensemble voting mechanism |
|
Cai Yujun1,2, Zhou Qiqing1,2, Guo Xi1,2, He Xiaolin3
|
|
1College of Land Resources and Environment,Jiangxi Agricultural University,Nanchang 330045,Jiangxi,China;2Key Laboratory of Agricultural Resources and Ecology in Poyang Lake Watershed,Ministry of Agriculture and Rural Affairs,Nanchang 330045,Jiangxi,China;3Jiangxi Provincial Agricultural Technology Promotion Center,Nanchang 330045,Jiangxi,China
|
| Abstract: |
| [Purpose] This study aims to construct a multi-dimensional collaborative data-cleaning framework that integrates statistical principles and professional mechanisms to enhance the reliability and research usability of cultivated land quality data.[Method] Taking the long-term time-series data from the four-level monitoring network(National-Provincial-Municipal-County)in Jiangxi Province as the foundation,a preliminary cleaning step based on domain-specific business logic constraints was introduced during the preprocessing stage. Subsequently,a three-dimensional parallel detection matrix was constructed,encompassing statistical distribution(3σ principle,IQR method),time-series features(Kalman filtering,moving average),and spatial structure(semivariogram,local Moran's I). An ensemble voting mechanism was then employed for cross-validation and comprehensive decision-making on the multi-dimensional identification results.[Result] (1)The framework exhibited good robustness in both soil attribute and crop yield data,effectively identifying multiple types of anomalies,including statistical,spatiotemporal,and logical errors. (2)The ensemble voting mechanism achieved a dynamic balance between anomaly identification and information preservation,successfully removing high-confidence anomalies while retaining genuinely high values,thereby overcoming the limitations of traditional single methods. (3)The introduction of domain knowledge in the preprocessing stage significantly improved the cleaning depth,using a logical verification mechanism to identify composite errors that statistical methods alone could not detect.[Conclusion] The proposed multi-dimensional collaborative cleaning framework shows significant advantages in achieving a dynamic balance between identifying anomalies and preserving essential information. It provides a reusable technical solution for the source governance and high-precision analysis of agricultural big data. |
| Key words: cultivated land quality data cleaning spatiotemporal heterogeneity ensemble voting mechanism |