predictor is disrupted. However, variable importance should be interpreted in the context of model validity
because explainability does not by itself demonstrate that a model has identified a robust or causal
environmental relationship. A predictor may receive a high importance score in a model that generalizes poorly,
particularly when the training data are limited, overfitted or unrepresentative of the intended prediction
domain.
[11,15]
For this reason, explainability should be considered alongside rigorous out-of-sample or cross-
domain validation rather than being treated as a substitute for predictive validity.
[11]
Comparing feature
rankings across held-out subbasins therefore provides a useful means of assessing whether the model identifies
relatively stable predictive relationships or whether the apparent importance of variables is specific to the
environmental conditions represented during model development. These distinctions have direct monitoring
implications. Machine learning may appear to reduce the measurement burden, but that claim is defensible only
if useful performance persists when the direct WQI components are unavailable and when the model is applied
outside its development domain.
[1]
A model that requires all index components may still assist rapid
computation or quality control, but it does not reduce the underlying measurement requirement. Similarly,
poor cross-subbasin performance limits claims of basin-wide deployment.
The present study is explicitly a methodological critique of machine-learning validation practices in hydrology
and is examined through a case study built around the Pra River Basin rather than a new assessment of current
water-quality conditions in the basin. It addresses this through a reproducible simulation informed by reported
characteristics of the basin.
[1]
A synthetic design, rather than the existing field dataset, was chosen for three
specific reasons. First, diagnosing target circularity and geographic transfer failure requires knowing the true
data-generating relationship between the predictors and the response; with observational field data, the true
functional form is unknown, so any transfer failure could be attributed either to a genuine methodological pitfall
or to unmodeled environmental heterogeneity, and the two explanations cannot be separated. A simulation in
which the WQI-generating rule is specified in advance removes this ambiguity: any degradation in cross-basin
performance can be attributed to the validation design itself rather than to unexplained field variability. Second,
a controlled benchmark allows the exact same underlying process to be regenerated under many independent
random draws (the 30-replicate Monte Carlo design reported below), which is not possible with a single fixed
field dataset and is necessary to show that the observed pattern is not an artifact of one particular sample. Third,
this design keeps the present study clearly distinct from, and complementary to, the existing field-based
geostatistical analysis of the same basin,
[1]
which already addresses spatial interpolation of observed WQI; the
present study instead isolates a purely methodological question—how validation design and predictor
definition affect apparent model performance—that a single observational dataset cannot answer in isolation.
This study therefore does not introduce new field observations; it builds a 150-record benchmark that
reproduces the previously reported WQI distributions and selected physicochemical characteristics for the
Birim, Offin and Pra subbasins,
[1,4]
and any conclusions drawn are conclusions about the validation
methodology, not about the present state of water quality in the basin. The full-feature and auxiliary-only
models are compared, the within-subbasin and leave-one-subbasin-out performance is evaluated, the stability
of feature importance is examined, a reduced-variable model is tested, whether the pattern is specific to tree-
based learners is tested by adding linear and ridge regression, and the experiment is repeated across multiple
independently generated datasets. The central objective is to determine whether strong local WQI prediction
remains convincing when the analysis is subjected to stricter tests of target circularity and geographic
generalization.
[1,4]
Accordingly, this study aims to evaluate the extent to which water quality index prediction is
influenced by the inclusion of the variables used directly in the construction of the index and to determine
whether the predictive relationships developed in the two subbasins can be transferred reliably to a third of
the previously unseen subbasin. The analysis further examines whether the observed patterns are robust
across model classes with fundamentally different extrapolation behaviors—the tree-based XGBoost and
random forest learners, which partition rather than extrapolate the training range, and linear regression and
ridge regression, which can in principle extrapolate linearly beyond it—to assess the stability of auxiliary
variable importance across subbasins using SHAP and permutation importance and investigate whether a
reduced set of auxiliary predictors can retain useful cross-subbasin predictive performance.
2. Materials and methods
2.1 Simulation framework and calibration targets
The analysis used a reproducible synthetic benchmark informed by summary statistics reported in the Pra
River Basin research program, specifically the 150-station monitoring network and subbasin WQI statistics
reported in Twenefour et al. (2026).
[1,4]
The benchmark contained 150 records, with 50 observations assigned
to each of the Birim, Offin and Pra subbasins. The synthetic WQI distributions were calibrated to the basin
means and standard deviations reported in that source, while the principal physicochemical variables were
calibrated to the corresponding subbasin means.
[1,4]
As set out in the Introduction, this design was chosen
specifically because it fixes the true data-generating relationship between predictors and WQI in advance,
which is what allows target circularity and geographic-transfer failure to be attributed unambiguously to