In brief: Instead of separately mapping water in pre- and post-flood images and then overlaying the results, the method first maps water features from optical and SAR imagery into a more consistent semantic space, then directly identifies where water conditions have changed. More importantly, it does not require large numbers of real flood-change labels: it can automatically construct pseudo-change samples from existing water-body datasets.
Flood emergency remote sensing faces a very practical dilemma.
Before a flood, clear optical imagery is often available. Afterward, however, heavy rain and thick cloud cover may persist, making SAR imagery—which can acquire data through clouds and in most weather conditions—the more reliable option.
As a result, the algorithm often receives not two visually similar images, but one optical image and one SAR image.
That creates a fundamental problem.
Optical imagery records reflected spectral information, while SAR measures microwave backscatter. The two look as though they speak entirely different languages. How can a change-detection model compare them directly?
In the paper Flood inundation monitoring using multi-source satellite imagery: a knowledge transfer strategy for heterogeneous image change detection, published in Remote Sensing of Environment, Wuhan University researchers Bofei Zhao, Haigang Sui and their colleagues proposed the Heterogeneous-Flood Inundation Extraction Network, or H-FIENet.
The work does not simply introduce a deeper neural network. It addresses two persistent challenges in emergency flood mapping at the same time: the mismatch between features in heterogeneous imagery and the shortage of real, bi-temporal flood-change samples.
1. Before Discussing the Network: Why Is “Pre-Flood Optical + Post-Flood SAR” So Difficult?

The Simplest Approach Is to Map Water Separately and Then Subtract the Results
A common method in conventional flood mapping is Post-Classification Comparison, or PCC.
Its logic is straightforward:
First, extract normal water bodies from the pre-flood image. Next, extract water from the post-flood image. Finally, compare the two binary water maps.
If a location was not water before the flood but becomes water afterward, it is classified as flood expansion. If it was water before but is no longer water afterward, it represents flood recession.
This sounds reasonable, but there is a major problem:
The accuracy of the final flood map depends on both the pre-flood and post-flood water-extraction results being correct.
Any misclassification in either image propagates into the final change map.
In SAR imagery, roads, building shadows and smooth land surfaces can all produce low backscatter and may therefore be mistaken for water. Errors from single-date water extraction are consequently carried into the flood map.
Why Not Let Deep Learning Compare the Two Images Directly?
That is precisely what change-detection networks are designed to do.
But this introduces a second problem: optical and SAR imagery are fundamentally different types of data.
Optical imagery depends on visible and near-infrared reflectance. SAR imagery depends on the interaction between microwaves and surface structure, roughness and dielectric properties.
The same river may appear as a dark, spectrally stable area in an optical image, while in SAR imagery it is primarily characterized by low backscatter.
The assumption underlying a conventional Siamese network—that its two branches can share weights—is therefore not naturally suited to this combination.
The Harder Problem Is the Shortage of High-Quality Flood-Change Labels
If hundreds of thousands of precisely co-registered pre-flood optical images, post-flood SAR images and pixel-level flood-change labels were available, a model could gradually learn the necessary relationships.
In reality, the opposite is true.
Large water-body extraction datasets are widely available, but datasets containing pixel-level flood-change labels across different satellites, regions and acquisition dates remain extremely limited.
Creating such samples manually is time-consuming, while flood emergencies themselves demand rapid response.
The paper therefore addresses two problems:
First, optical and SAR imagery occupy different feature spaces.
Second, heterogeneous flood-change samples are severely limited.
H-FIENet was developed around these two challenges.
2. H-FIENet’s Core Idea: Do Not Learn “Flooding” From Scratch—Transfer Existing Knowledge About Water

The authors identified an important fact:
Although heterogeneous flood-change data are scarce, single-date water-body extraction data are not.
Rather than asking a new model to learn from scratch what water is, what water looks like in optical imagery, what it looks like in SAR imagery and when a change should be classified as flooding, the method first inherits basic water knowledge from pretrained water-extraction models.
The overall logic can be summarized in one sentence:
First, teach the optical and SAR branches to recognize water independently. Then let the change-detection module learn what has happened to the water.
The paper calls this design cross-task knowledge transfer.
Water-body extraction is the source task, while flood-change detection is the target task.
H-FIENet therefore does not simply transfer a conventional classification model to a flood application. It transfers the water-related semantic knowledge most relevant to flood detection.
3. The First Key Innovation: Optical and SAR Each Speak Their Own Language


A conventional Siamese network usually consists of two branches with identical structures and shared parameters.
This is reasonable for change detection using homogeneous imagery because the pre- and post-event images come from the same or similar sensors and generally have consistent feature representations.
For optical and SAR imagery, however, sharing parameters can become a constraint.
H-FIENet instead adopts a pseudo-Siamese architecture.
The two branches do not share parameters. Each uses a water-feature extraction network suited to its own data type.
In the paper:
- The optical imagery branch uses MSResNet.
- The SAR imagery branch uses Siam-DWENet.
The two models are first trained separately on optical and SAR water-body datasets.
The authors then remove the output components originally used for single-date water classification while retaining the intermediate ability to extract water-related features. These become H-FIENet’s two feature transformation modules.
The process therefore becomes:

What matters here is not simply that the network has two branches.
Instead of directly comparing raw optical and SAR pixels, the branches first transform their inputs into more comparable water-related semantic representations.
In other words, the model does not have to answer:
“Is this SAR intensity value similar to that optical pixel value?”
It instead asks:
“After sensor-specific water-feature extraction, has the water state represented at this location changed?”
This is the core mechanism H-FIENet uses to address the feature mismatch between heterogeneous images.
4. Why Is Feature Alignment More Sensible Than Direct Image Translation?

One common approach to optical–SAR change detection is to “translate” one image type into the other and then calculate the difference.
For example, a model might attempt to regress pre-flood optical features into a representation resembling post-flood SAR features before comparing them.
But this creates another problem.
Images contain extensive information about buildings, roads, farmland, vegetation, water and other land-cover types. Flood detection, however, is concerned primarily with whether the water state has changed.
H-FIENet therefore does not require the two modalities to become fully consistent across every type of land-cover feature. Instead, it focuses on aligning deep semantic features related to water.
Figure 12 in the paper compares the feature transformation results produced by AGSCC, CAAE and H-FIENet.
AGSCC can reveal prominent change areas through feature regression, but its superpixel processing introduces jagged boundaries and loses detail. CAAE relies on unchanged areas for unsupervised feature learning, making it prone to missed detections when changed areas account for a large proportion of the image.
H-FIENet instead concentrates on water-related semantics.
Its goal is therefore not:
To make optical imagery look like SAR imagery.
It is:
To make “water” in optical imagery and “water” in SAR imagery easier to compare within a deep feature space.
That is the most important point for understanding H-FIENet.
5. The Second Key Innovation: Creating Change Labels Without Real Paired Flood Samples


Resolving the feature mismatch is not enough. Deep-learning-based change detection still typically requires large volumes of labeled bi-temporal data.
One of the paper’s most instructive ideas is to reconsider a basic assumption:
Must the pre-event and post-event images used for change-detection training come from the same location?
The authors’ answer is no.
What the Model Really Needs to Learn Is a Change Between Water and Non-Water
Suppose an optical image from China already has a water mask.
A SAR image from another region also has a water mask.
Conventional change-detection methods would not pair them because they depict entirely different locations.
The authors argue, however, that geographic correspondence is unnecessary if the sole objective is to teach the model that a transition from water to non-water—or from non-water to water—represents change.
Optical water samples and SAR water samples from different regions can therefore be paired at random.
Only their water labels are then compared.
If the two labels indicate different states, the pixel is marked as “changed.” If they indicate the same state, it is marked as “unchanged.”
The logic is straightforward:
In essence, the method applies an exclusive-or operation to the two water masks.
The result is a pseudo flood dataset generated from spatially inconsistent multi-source pre- and post-event imagery.
Why Does This Matter?
Because it changes how training samples can be obtained.
The conventional route is:
Find pre- and post-event images of the same location → precisely co-register them → manually label flood-related changes
H-FIENet instead uses:
Existing optical water data + existing SAR water data → random pairing → automatic generation of pseudo-change labels
Flood-change labels are scarce, but water masks are widely available.
The authors therefore turn a difficult data-acquisition problem into one that can reuse existing public datasets.
6. Can the Model Really Be Trained With Just Nine Sample Pairs? Transfer Learning—not the Number Alone—Makes It Possible
One number in the paper immediately stands out:
H-FIENet uses only nine sets of 512 × 512 image patches for flood-change training.
Each set contains a pre-event optical image, a post-event SAR image and an automatically generated pseudo-flood-change label. The nine sets are divided among training, testing and validation at a ratio of 4:4:1.
A conventional deep network initialized with random weights would be unlikely to learn complex water representations across different satellite sensors reliably from only nine change-detection samples.
H-FIENet, however, does not start from scratch.
Its two front-end branches have already been pretrained on optical and SAR water datasets, allowing them to acquire substantial knowledge about water features.
The final few-shot training stage is therefore less about teaching the model what water is and more about teaching it how to compare water features that have already been identified.
This is why the paper’s two main elements must be considered together:
Cross-task knowledge transfer + pseudo-change sample generation
Pseudo-samples alone would not be enough to support a robust model without stable, pretrained water features. Pretrained features alone would not eliminate the training-data bottleneck without a low-cost way to generate change labels.
Together, the two components form a complete strategy.
7. How Was the Method Tested? Across Different Countries and Satellite Combinations
To show that H-FIENet was not effective only for one pair of Sentinel images, the authors evaluated it across multiple flood scenarios and satellite combinations.
The principal datasets included:
The study areas included flood-affected regions around Wangjiaba in China, Liège in Belgium and Maputo in Mozambique.
These areas differ substantially in land cover, encompassing farmland, settlements, rivers, wetlands and vegetation across different climatic and geographic environments.
The experiments were therefore designed to answer more than whether the model could accurately map one flood.
The broader question was:
Would the feature-transfer strategy continue to work in another country, over different terrain or with a different satellite combination?
8. How Did It Perform? Overall Accuracy Reached 0.943 at Wangjiaba

For the Wangjiaba experiment, the authors used pre-event Sentinel-2 optical imagery and post-event Sentinel-1 SAR imagery to identify flood inundation.
H-FIENet was compared with CMCDNet, GIR-MRF, AGSCC, CAAE and INLPG.
The main results were:
The most important point is not simply that 0.943 is higher than 0.930.
CMCDNet is a supervised heterogeneous change-detection method that requires a larger number of bi-temporal training samples. H-FIENet achieved comparable or better results using only a small number of pseudo-change samples.
The paper also reports that H-FIENet processed the 3,800 × 3,100-pixel pre- and post-event images in approximately one minute.
That processing speed brings the method closer to the rapid-response requirements of emergency flood mapping.
9. Does It Still Work in Europe and Africa?


The authors conducted further tests in Belgium and Mozambique, using flood products from the Copernicus Emergency Management Service as reference data.
H-FIENet achieved the following overall accuracy scores across the different scenarios:
- EMSR518: 0.954
- EMSR650-02: 0.871
- EMSR650-03: 0.933
In the EMSR650-02 scenario, H-FIENet did not achieve the highest OA among all the methods. Its Kappa score nevertheless reached 0.441, and it produced a comparatively strong overall balance between false alarms and missed detections.
These experiments matter because they show that the model’s effectiveness was not confined to the Wangjiaba case. It continued to provide usable change-detection results under very different flood conditions in Europe and Africa.
In the Mozambique experiment, the paper also reports that H-FIENet mapped the flood extent within several minutes and could execute the detection process automatically, whereas the CEMS products involved expert input and post-processing.
The emphasis, therefore, is on the method’s potential for automated emergency mapping—not merely its pixel-level accuracy under experimental conditions.
10. Another Satellite Pair: Gaofen-2 and Gaofen-3

If a model can process only Sentinel-2 and Sentinel-1 imagery, its claim to support “multiple sources” remains limited.
The authors therefore tested H-FIENet on the Wangjiaba flood using:
Pre-flood Gaofen-2 optical imagery + post-flood Gaofen-3 SAR imagery
The results showed that H-FIENet could still delineate the main flooded areas.
This experiment validates an important principle behind the H-FIENet architecture:
The change-detection stage can remain relatively independent of the sensor-specific water-feature modules.
If a new remote-sensing data source is introduced, the entire change-detection framework may not need to be rebuilt. Instead, a corresponding water-feature transformation module could be developed for the new sensor and integrated into H-FIENet.
This modular design offers greater potential for expansion than a network tied to a fixed sensor pairing.
11. An Easily Overlooked Strength: It Detects Both Flood Expansion and Recession
Many flood studies ultimately answer only one question:
What was the maximum extent of the flooding?
But floods are dynamic. Water expands and then recedes.
The authors therefore used multi-temporal imagery from Wangjiaba acquired on July 20, August 1 and August 17, 2020, to test whether H-FIENet could process different image combinations.

SAR → SAR: Detecting Both Expansion and Recession
Using Sentinel-1 SAR imagery from July 20 and August 1, H-FIENet identified both newly inundated areas and areas where floodwaters had receded.
The results were:
- Flood-expansion detection OA: 0.973
- Flood-recession detection OA: 0.853
For flood expansion, H-FIENet performed similarly to the conventional RI and NCI methods. For flood recession, the paper found that H-FIENet showed a clearer advantage in its spatial results.

SAR → Optical: It Also Works When the Data Order Is Reversed
The authors then used SAR imagery from August 1 and optical imagery from August 17 for change detection.
This time, the input sequence was no longer “optical → SAR,” but:
SAR → optical
H-FIENet was still able to identify areas where floodwaters had receded.
This suggests that the model is not intended solely for a fixed “pre-flood optical, post-flood SAR” configuration. To some extent, it supports more flexible combinations of bi-temporal remote-sensing data.
The study therefore takes a step beyond one-off emergency flood mapping toward multi-temporal monitoring of flood evolution.
12. The Real Innovation Is More Than a Pseudo-Siamese Network
Judging only by the architecture, it would be easy to summarize this study as:
“The authors designed a dual-branch change-detection network.”
But that would understate its contribution.
What makes the study notable is how it reorganizes the relationship among data, knowledge and models in flood-change detection.
The study’s contribution can therefore be distilled into three concepts.
Cross-Task Knowledge Transfer
Knowledge learned from water extraction is transferred to flood-change detection, reducing dependence on flood-change labels.
Heterogeneous Feature Alignment
Rather than forcing optical and SAR imagery through the same feature extractor, the two modalities are processed separately and compared in a higher-level water-semantic space.
Pseudo-Change Samples
The pre- and post-event images no longer need to come from the same location. Existing water labels can instead be used to generate change supervision automatically.
Together, these three elements form H-FIENet. They are not isolated techniques.
13. What Does This Approach Mean for Emergency Flood Monitoring?
No Need to Wait for the “Ideal” Post-Disaster Data Combination
After a flood, there is no certainty about which satellite will provide the first usable imagery.
If an algorithm accepts only a fixed Sentinel-2 → Sentinel-1 combination, the emergency response system remains constrained by data availability.
H-FIENet aims to make use of whichever suitable data source becomes available first.
As long as a usable water-feature extraction module exists for the relevant sensor, its imagery could potentially be incorporated into the same change-detection framework.
Shifting the Data Bottleneck From Flood-Change Labels to Water Labels
Flood-change labels are expensive and scarce, while large volumes of single-date water data already exist worldwide.
One of H-FIENet’s main contributions is that it allows these existing datasets to support not only water extraction but also flood-change detection.
Easier Adaptation to New Regions and Satellites
Because the two sensor branches are relatively independent, integrating a new optical or SAR sensor could begin with training a corresponding water-feature module.
This reduces the need to rebuild the entire flood-change model whenever a new satellite is introduced.
A Foundation for Time-Series Flood Monitoring
The paper demonstrates that H-FIENet can process SAR–SAR and SAR–optical combinations and distinguish between flood expansion and recession.
This gives the approach potential for organizing multi-temporal imagery into dynamic flood sequences.
14. Conclusion: What Matters Most Is Not 0.943, but “Align Water Semantics Before Detecting Change”
H-FIENet achieved an OA of 0.943 in the Wangjiaba heterogeneous optical–SAR flood experiment. It also maintained solid detection performance across flood scenarios in Europe and Africa, extended to Gaofen-2/Gaofen-3 imagery, and identified both flood expansion and recession.
More important than these figures, however, is how the study breaks down the problem.
The conventional approach is often to make the model learn all the complex differences between optical and SAR imagery directly.
H-FIENet reframes the problem:
Optical and SAR imagery may look completely different, but both can first answer the question, “Is there water here?” Once their water-related semantic representations are aligned, the remaining task is simply to determine whether the water state has changed.
The authors also convert scarce bi-temporal flood-change labels into pseudo-change samples that can be generated automatically from existing single-date water labels.
The study therefore addresses more than the design of an individual network layer. It explores:
How existing knowledge and data can help heterogeneous remote-sensing imagery from multiple sources produce comparable, transferable and more automated change results more quickly during flood emergencies.
That is what makes the study particularly noteworthy following its publication in Remote Sensing of Environment.
It moves flood-change detection beyond the assumption that ideal bi-temporal training data must always be available.
Even when imagery comes from different satellites and modalities, and genuine flood-change samples are scarce, knowledge transfer can make existing water datasets useful again.
Organizations planning operational flood-monitoring workflows based on optical and SAR imagery can contact STARPATH GLOBAL to discuss coverage, revisit frequency and processing requirements. China’s expanding satellite capacity gives international customers access to competitively priced imagery, while our imagery catalog helps match resolution to the actual use case so teams do not pay for more detail than they need. Organizations without in-house remote-sensing experience can also apply to the Pioneer Partner Program, where our FDE team works directly with customer teams to evaluate use cases, workflows and ROI before commercial deployment.
Paper Information
Title: Flood inundation monitoring using multi-source satellite imagery: a knowledge transfer strategy for heterogeneous image change detection
Authors: Bofei Zhao, Haigang Sui, Junyi Liu, Weiyue Shi, Wentao Wang, Chuan Xu and Jindi Wang
Journal: Remote Sensing of Environment
Volume and article number: Volume 314, 114373
Year: 2024
DOI: 10.1016/j.rse.2024.114373





