-1000% occupancy: how I found holes in my own published dataset
The author shares their experience of discovering critical errors in a published open dataset regarding parking occupancy. After an anomalous occupancy value of -1000% was recorded, an investigation began, revealing serious data quality issues. During the process, the author found that standard recalculation methods were ineffective, and finding the missing denominator required a non-standard approach. The article details why automated checks, which had been running for a year and a half, failed to detect the anomaly in time. This case serves as an important reminder of the need for thorough data validation before publication and highlights that even tested systems can harbor hidden defects. The author analyzes their mistakes and offers lessons to help other developers and analysts avoid similar situations when working with large datasets.
This is a summary. Read the full article at the original source:
HabrRelated stories
This Habr article explores the methodology of switchback experiments, which are essential for evaluating changes in systems with shared resources wher…
I built a self-hostable AI data analyst — 4 agents, sandboxed code execution, your choice of LLM
Developer Laban has released Insight Orchestra, an open-source, self-hostable AI data analysis platform designed to replace cloud-based alternatives l…
In this Habr article, the author explores a fundamental issue in statistical inference: how the experiment stopping rule affects p-values. Using a coi…



