Abstract
Outliers significantly undermine the robustness of regression analysis. However, classical robust regression methods often fail to effectively handle diverse data scenarios. To address this limitation, we propose a novel outlier-detection framework that fundamentally differs from conventional approaches by leveraging data morphological features and residual structural characteristics to ensure robust performance across various outlier types, sample sizes, dimensions, and contamination rates. This framework operates through three sequential stages: partition, partial purification, and full recovery. The partition stage employs a recursive clustering-based approach to segment data into spatially adjacent and linearly structured components, facilitating outlier isolation. The partial purification stage then extracts a reliable subset of clean samples through a strategy based on least median of squares (LMS), providing an initial robust estimate. Finally, the full recovery stage identifies clustered and sparse outliers by examining the shape and magnitude of robust residuals, enabling comprehensive clean data reconstruction. Extensive evaluations on synthetic and real-world datasets confirm the superior performance of the method across challenging conditions, as validated through statistical analysis, sensitivity analysis, and scalability analysis.
| Original language | English |
|---|---|
| Article number | 115187 |
| Pages (from-to) | 1-15 |
| Number of pages | 15 |
| Journal | Knowledge-Based Systems |
| Volume | 335 |
| DOIs | |
| Publication status | Published - Feb 2026 |
Fingerprint
Dive into the research topics of 'A shape-enhanced outlier-detection framework for regression tasks'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver