A Hybrid Bias Correction Framework for Daily Satellite Precipitation
Date
2026Jenis/Type
TesisSubtype
ThesesAuthor
Istanto, Benny
Boer, Rizaldi
Santikayasa, I Putu
Metadata
Show full item recordAbstract
Daily satellite precipitation underpins flood early warning, drought monitoring, and water-budget analysis, but the Integrated Multi-satellitE Retrievals for Global Precipitation Measurement Late Run (IMERG-L) carries systematic biases over the Indonesian Maritime Continent. This thesis evaluates a four-stage hybrid bias correction pipeline: Linear Scaling (LS), Empirical Quantile Mapping (EQM) with a Gamma fit, Generalized Pareto Distribution (GPD) tail adjustment, and a station-density-aware Convolutional Neural Network (CNN) refinement, a Deep Learning (DL) architecture, denoted LSEQM+DL. The pipeline is applied to IMERG-L over Indonesia across 2001 to 2025, with gauge validation against 172 out-of-sample stations of Indonesia’s Meteorological, Climatological, and Geophysical Agency (BMKG) over 2001 to 2021.
Against the out-of-sample BMKG network, three of the four verification pillars move close to the gauge target. Relative bias drops from -11.4% for the raw Linear-Scaling stage to -0.6% for LSEQM+DL, the standard-deviation ratio moves from 0.71 to 1.00, the Kolmogorov-Smirnov p-value rises from 0.01% to 19.1%, and the ratios at the 95th and 99th percentiles move from 0.74 and 0.71 to 1.05 and 1.01. Event detection shows a designed trade-off: the full pipeline detects fewer light-rain events than the Linear-Scaling stage at and below 10 mm/day but achieves higher critical success at every threshold from 20 mm/day upward, with a 45% Critical Success Index gain at the 50 mm/day threshold (0.091 versus 0.062). The crossover coincides with the internationally standardised very-heavy-precipitationday threshold of 20 mm/day, the regime in which flood and drought decisions are made.
One diagnostic barely moves. Measured per pixel against the in-sample CPCUNI target, the Pearson correlation between paired daily values stays near 0.35 across all three corrected stages. It also stays in the narrow band [0.332,0.348] across all fifteen settings of a Bali sensitivity sweep over the blending weight, the GPD threshold percentile, and the station-density saturation count, none of which was formally optimised. This is the predicted behaviour of marginal bias correction, where rank preservation and variance inflation without time-dependent skill leave the correlation pinned close to that of the raw retrieval.
A separate and distinct effect governs the comparison against the independent BMKG stations, whose daily totals close at the morning observation and carry a date label one day later than the UTC-dated satellite and CPC-UNI archives. The satellite can be matched to that label in two ways: by shifting the date of the daily archive by one day, or by re-aggregating the half-hourly source to a -23-hour window. Either raises the pooled correlation over the GPM era (2015 to 2021) from r = 0.20 to r = 0.57, with a per-station median of 0.56. The daily archive alone recovers almost all of the gain, so the half-hourly source is not required. The CPC-UNI calibration correlation carries no such offset and is not improved by the shift. Nor does the correction itself improve the day-by-day timing once the labels are matched, which places the remaining ceiling in the timing of the raw retrieval rather than in the calendar alone.
The same -23-hour offset is optimal in all twelve three-month running seasons and all three Indonesian time zones, while the recovered band correlation ranges from 0.55 in the western and central zones (WIB and WITA) to 0.61 in the eastern zone (WIT). The recovery is confined to the GPM era, in which the retrieval resolves the diurnal cycle, so the full-record correlation against BMKG blends this era with the earlier, timing-imprecise era. The window dependence dissolves under monthly aggregation, where the correlation is near 0.80 regardless of window, placing the daily ceiling in the timing of the day-by-day pairing rather than in the rainfall the product captures.
This date-label matching is a single-stage pre-processing change outside the correction pipeline, recoverable from the daily archive alone, and is identified as the most operationally actionable next step. The corrected product is suitable as delivered for applications that depend on the daily distribution, the regime the reaggregation leaves unchanged.
The pipeline is released as an open, configuration-driven implementation, and its reproducibility is demonstrated by re-executing the headline results on free cloud infrastructure. The correction, verification, validation and visualisation sequence over the Bali subdomain (9×14 grid at 0.1º, 80 land pixels, all 36 dekadal windows of the 2001 to 2025 record) completes in 72.1 minutes on a standard Google Colab Central Processing Unit runtime.

