construction, and labeled data is scarce. Interpretable classical machine learning and unsupervised
autoencoders, deployed passively, are the realistic choice. Edge and resource-constrained deployments, such as
remote substations and frames, rule out GPU-class inference and favor compact classical models or quantized
one-dimensional convolutional networks. Safety-critical, SIL-rated loops cannot tolerate obscured or actuating
models. Interpretable detection in advisory mode is the only currently defensible option, and autonomous RL
is contraindicated until safe-RL and verification methods mature, as Section 7.2 discusses. Finally, large and
dynamic enterprise-OT networks, where lateral movement is the dominant concern, are where topology-aware
GNNs justify their construction and maintenance overhead.
6.4 Dataset bias and generalization
The resulting bias deserves explicit treatment rather than a passing caveat. Among the 102 quantitatively
assessed studies, a large majority evaluate on one of four public corpora: SWaT, WADI, the Mississippi State gas-
pipeline and power-system datasets, and BATADAL. The first is benchmark overfitting at the level of the field
rather than the individual model. When consecutive studies tune architectures against the same fixed masses,
reported developments partly reflect adaptation to that corpus, including its specific attack scripts, sensor
configuration, and noise characteristics. Gains of one or two percentage points on SWaT, accumulated across
many papers, cannot be assumed to transfer to an unseen plant.
The second bias deals with limited attack realism. Public ICS datasets are created by executing a predetermined
and usually modest catalogue of attack scenarios against a testbed. Attacks are typically launched by researchers
rather than by adversaries who adapt to the defense, and stealthy long-horizon campaigns typically seen in
TRITON are usually under-represented. A detector validated on such data has been tested against a fixed and
comparatively cooperative opponent.
The third isuue is scale and sector skew. SWaT and WADI are small water-treatment testbeds with tens of
sensing setpoints. Production ICS environment routinely have thousands, with heterogeneous vendors and
partially documented topology. Water treatment and power are heavily over-represented relative to oil and gas,
chemical processing, and discrete manufacturing. Verification for sector transfer is consequently thin, which is
why Section 5.5 treats apparent cross-sector consistency with caution.
The fourth is class imbalance and metric fragility. Attack samples form a small minority of most corpora.
Accuracy is therefore an unreliable summary, and a detector predicts the majority class can score highly. This is
the reason Table 6 records false-positive rates separately and marks unreported values rather than imputing
them.
Two mitigations are visible in the more careful recent studies and are worth adopting as norms. To address this,
cross-dataset evaluation in which a model trained on one corpus is tested on another without retuning is
imperatively recommended. This study infirmly substantiate accuracy degradation as observed in the few
studies. Secondly, the sensor configuration, sampling rate, and process modes under which a result holds should
be explicitly reported within the operational envelope. Neither practice is yet common, and their absence is
recorded as a methodological gap.
6.5 Strengths and limitations synthesis
Table 8 consolidates the comparative evidence into a strengths-and-limitations matrix as shown in scenario
mapping of Table 7. No single family satisfies the joint requirements of accuracy, latency, interpretability,
topology awareness, and autonomous response. Layered architectures represent the most credible deployment
pattern. Such architectures use classical machine learning for fast triage, deep models for high-fidelity detection,
GNNs for network-wide risk propagation, and eventually constrained RL for response orchestration. This
mirrors the defense-in-depth philosophy already embedded in IEC 62443 zone-and-conduit design.
[29,30]
6.6 Identified research gaps
The four methodological gap clusters revolve across the corpus. These methodological gaps include the absence
of standardized benchmarks and metrics, which prevents meaningful cross-study comparison; over-reliance on
a small number of testbed datasets, with the consequences set out in Section 6.4; and adversarial robustness,
which is examined in fewer than 10% of studies despite its evident relevance to a contested environment.
[34,44]
Application gaps include millisecond-level latency validation in production settings, retrofitting onto legacy and
resource-constrained devices, and safetysecurity co-engineering under IEC 61508 and IEC 62443.
[29,30]
Technical gaps include multi-stage kill-chain correlation, as most detectors flag isolated events; transfer
learning across sectors and protocols; and humanAI collaboration, covering trust calibration, alarm
rationalization, and operator interface design.
[13,44]
Data and privacy gaps include attack-data scarcity, barriers
to cross-facility sharing, and non-independent and non-identically distributed data across sites. Table 9 maps
each cluster to the directions developed in Section 7.