The question to address for this stage:
Can a subtle signal be embedded directly into an image, survive ordinary image processing, and still be detected blindly without the original image, metadata or an external database?
The purpose of this stage was to determine whether the underlying image-native communication mechanism was technically plausible. By the end of Stage 1, we had built the first complete GAPP research pipeline:
Original image
↓
Content-adaptive encoder
↓
Multi-carrier pixel signal
↓
Protected image copy
↓
Image processing
↓
Blind decoder
↓
NOT_PROTECTED / INDETERMINATE / PROTECTEDThe system worked well enough to justify continuing the research, but Stage 1 also revealed a major architectural limitation: cropping and reframing disrupted the global coordinate system on which the signal depended.
Starting point
Stage 1 began with almost nothing. There was no trained detector, no signal architecture and no established evaluation methodology. The earliest experiments were deliberately simple: add a known perturbation to an image and determine whether that perturbation could still be detected after the image had been processed. The initial goals were therefore narrow:
- Embed information directly into image pixels.
- Keep the visual change small.
- Detect the signal using only the received image.
- Survive common image operations.
- Avoid classifying ordinary images as protected.
- Establish an evaluation process that did not repeatedly tune against the same test images.
The architecture became more sophisticated only when the experiments showed why the simpler approaches failed.
What we explored
1. The signal evolved from random noise to structured carriers
The first experimental signal used random per-pixel perturbations. That was useful for proving the basic encode–transform–detect pipeline, but it had an obvious weakness: resizing averages neighbouring pixels together. High-frequency random structure was therefore damaged heavily when an image was resized. JPEG compression was comparatively survivable. Resizing was not. That led to the first major design change.
Random per-pixel signal
↓
Low-frequency signal
↓
Content-adaptive low-frequency signal
↓
Three independent RGB carriersInstead of distributing unrelated perturbations across individual pixels, the
signal was moved into a lower-frequency spatial structure. This made the signal
substantially more stable under resizing and compression. Later experiments
introduced three independent carriers, one associated with each RGB channel.
The intention was not to make the signal stronger. It was to make accidental
alignment with ordinary image content less likely. At the same empirical 1%
false-positive operating point, the raw multi-carrier comparison improved protected
detection from:
Single carrier: 11.20%
Multi-carrier: 41.50%while the mean pixel change actually fell slightly:
Single carrier: 0.01347
Multi-carrier: 0.01295Ordinary-image score variation was reduced by approximately 50.9%. That became an
important Stage 1 result: multiple independent carriers provided cleaner statistical
evidence without requiring a stronger visible perturbation.
2. The encoder learned that images have different signal capacity
Applying exactly the same signal everywhere was also too simplistic. Different image regions tolerate perturbations differently. A textured or visually complex region may conceal a small signal easily, while smooth gradients, flat colours or visually sensitive regions can make the same perturbation more noticeable. The encoder therefore became content-adaptive. Rather than applying uniform strength across the image, it estimated where signal could be placed more safely and adjusted the embedding accordingly. This also introduced the concept of image capacity: an estimate of how much usable signal an image could carry while preserving its appearance. Capacity became one of the strongest predictors of detection reliability. For example, in the later multi-carrier system:
Capacity above 0.40
Clean detection: ~97%
Resize + JPEG detection: ~97%
Capacity above 0.50
Clean detection: 100%
Resize + JPEG detection: 100%Low-capacity images remained much harder. We tested whether stronger encoding could compensate for this by increasing signal strength on low-capacity images. It sometimes improved individual cases, but it also risked introducing visible texture and did not reliably improve the overall operating point. The conclusion was therefore not simply "stronger signal = better system". Rather, Stage 1 showed that image content places a real constraint on how much signal can be embedded cleanly.
3. The decoder could not behave like an ordinary image classifier
The first learned decoder experiments were also instructive. A conventional CNN trained directly on protected and ordinary images did not naturally discover the subtle signal we wanted it to detect. It tended to remain near chance, overfit, or learn properties of the image content rather than the carrier itself. The decoder became more useful once we explicitly exposed the structure we already knew existed. Its evolution was roughly:
Direct correlation
↓
Raw CNN
↓
Pattern-aware decoder
↓
Capacity-aware decoder
↓
Pairwise-trained capacity-aware decoder
↓
Multi-carrier capacity-aware decoderThe final Stage 1 decoder used evidence associated with the known carriers together with image context and capacity-related features. Training also changed. Because every protected image had a matching ordinary version, the model could be taught not only to classify images independently but also to learn that:
score(protected copy) > score(ordinary copy)This pairwise structure became more useful than treating every sample as unrelated.
What we built
The frozen Stage 1 candidate used:
| Component | Final Stage 1 configuration |
|---|---|
| Signal | Low-frequency multi-carrier |
| Carriers | Three independent RGB patterns |
| Grid | 16 Ă— 16 |
| Signal strength | 0.08 |
| Encoder | Content-adaptive |
| Decoder | Blind, capacity-aware, carrier-aware |
| Training | Pairwise classification + ranking |
| Original image required | No |
| Metadata required | No |
| External database required | No |
The three carrier seeds were:
2026
2027
2028An early research tri-state decision rule was also introduced:
if score ≤ 0.08713963 = NOT_PROTECTED
if 0.08713963 < score < 0.96602160 = INDETERMINATE
if score ≥ 0.96602160 = PROTECTEDThese Stage 1 thresholds were experimental detector thresholds. The deeper protocol meaning of the three states was not formally defined here.
Key results
Once the architecture was selected, we stopped modifying it and evaluated the frozen system on increasingly separated data.
Development set
The multi-carrier learned decoder reached:
| Condition | Protected detection | False positives |
|---|---|---|
| Clean | 78.30% | 0.60% |
| Resize 50% + JPEG 75 | 76.20% | 0.70% |
This was substantially better than the preceding single-carrier pairwise system.
Untouched Caltech reserve
The next evaluation used all 3,427 images in the previously
untouched Caltech reserve exactly once.
| Condition | Protected detection | False positives | Determinate accuracy |
|---|---|---|---|
| Clean | 81.41% | 0.85% | 97.93% |
| Resize 50% + JPEG 75 | 79.14% | 0.96% | 97.61% |
The reserve did not reveal a major collapse. This provided evidence that the development result was not simply caused by one favourable split.
Oxford-IIIT Pet
The frozen model and thresholds were then applied to a completely different
dataset with no retraining or recalibration. 3,669 Oxford-IIIT Pet test images
were evaluated.
| Condition | Protected detection | False positives | Determinate accuracy |
|---|---|---|---|
| Clean | 92.56% | 0.57% | 99.49% |
| Resize 50% + JPEG 75 | 91.47% | 0.52% | 99.50% |
Oxford performed considerably better than Caltech. Later capacity analysis showed why: Oxford contained far fewer low-capacity images. Its stronger result therefore provided useful cross-dataset evidence, but it should not be interpreted as a universal expected detection rate.
Pascal VOC 2007
The same frozen system was then evaluated on 4,952 Pascal VOC 2007 test images.
| Condition | Protected detection | False positives | Determinate accuracy |
|---|---|---|---|
| Clean | 90.06% | 0.97% | 99.17% |
| Resize 50% + JPEG 75 | 88.59% | 0.99% | 99.06% |
Across Oxford and VOC together, the frozen system therefore produced approximately:
Clean
Protected detection: 91.13%
False positives: 0.80%
Resize 50% + JPEG 75
Protected detection: 89.82%
False positives: 0.79%Note: These figures are evidence from the tested datasets only. They do not represent a claim of universal performance.
What worked
Several Stage 1 ideas survived later scrutiny.
-
Low-frequency structure survived ordinary processing: Moving away from random high-frequency perturbations made the signal substantially more resistant to frame-preserving resizing and JPEG compression.
-
Multiple carriers reduced accidental matches: The RGB multi-carrier design lowered ordinary-image score variability and produced a major improvement at low false-positive operating points.
-
Content-adaptive encoding was useful: Images do not offer equal embedding capacity. Allowing the encoder to respond to image structure worked better than treating every image identically.
-
Explicit signal knowledge helped the decoder: The strongest decoder did not discover everything from raw pixels alone. Giving it structured carrier evidence proved much more effective than relying on an unconstrained CNN.
-
Pairwise training matched the actual problem: Protected and ordinary copies naturally form pairs. Learning their relative ordering produced useful information that independent classification objectives missed.
-
Frozen cross-dataset evaluation transferred: The selected system retained strong performance on Oxford-IIIT Pet and Pascal VOC without retraining or recalibration.
What did not work
Failed approaches were kept as part of the research record because they shaped the architecture.
-
Random per-pixel signalling: It survived JPEG reasonably well but deteriorated badly after resizing. Lesson: high-frequency random structure was too fragile.
-
Reconstructing the encoder's adaptive mask inside the decoder: The decoder could not reliably reproduce the exact adaptive mask after the image had already been modified by encoding and processing. Lesson: encoder-side adaptation cannot simply be assumed recoverable by the detector.
-
Raw CNN decoding: A conventional image classifier struggled to discover the subtle known carrier and tended to overfit or learn image content. Lesson: the decoder needed structured signal-aware evidence.
-
Capacity-normalised strength: Increasing signal strength on difficult images helped some cases but could reintroduce visible artefacts and did not cleanly solve low-capacity detection. Lesson: low image capacity is not solved simply by increasing amplitude.
-
Hard-negative weighting: Making training more conservative around difficult ordinary images reduced some false positives but also damaged protected detection substantially. Lesson: false-positive reduction cannot come at unlimited recall cost.
-
Early model ensembling: Combining similar single-carrier models did not produce meaningful gains because they often made the same mistakes. Lesson: an ensemble only helps when its members contribute genuinely different evidence.
-
Capacity-weighted positive training: Giving additional weight to low-capacity protected examples did not specifically rescue them. Lesson: the main problem was not simply insufficient optimisation pressure on difficult examples.
Cropping exposed the architectural limit
The strongest Stage 1 failure appeared only after the clean and resize/JPEG system was already working well. The frozen system was tested under cropping and reframing. Performance collapsed, particularly for off-centre crops. The reason was increasingly clear: the Stage 1 carrier was defined relative to the complete image coordinate frame.
Full image
↓
Known carrier geometry
↓
Strong detection
Crop or reframe
↓
Scale + position change
↓
Carrier geometry changes
↓
Detection fallsAn oracle experiment tested whether knowing the exact crop geometry could restore
the result. For off-centre 90% crops, alignment clearly helped:
Top-left crop
0.53% → 14.01% protected detection
Bottom-right crop
5.11% → 16.18% protected detectionBut even exact alignment recovered nowhere near the roughly 90% clean result. This
rejected the simple explanation that cropping was only an alignment problem. The
learned decoder itself had become specialised around the original global carrier
geometry and evidence distribution. That became the limitation for this stage.
What Stage 1 taught us
Stage 1 answered its original question (defined at the beginning) positively. But it also changed the research question. The main problem was no longer:
Can a signal survive?
It became:
Can protection evidence survive locally when much of the original image and its coordinate system are gone?
The crop experiments suggested that a future architecture should behave less like one watermark spread across an entire frame and more like many redundant local pieces of evidence. A sufficiently large surviving region should ideally contain enough information to support detection on its own.
Outcome
STAGE 1
COMPLETE
Established:
âś“ Image-native signalling is technically plausible
âś“ Blind detection is possible
âś“ JPEG robustness
âś“ Frame-preserving resize robustness
âś“ Cross-dataset transfer
âś“ Low false-positive operation
âś“ Content-adaptive encoding
âś“ Multi-carrier signalling
Unresolved:
â–ł Low-capacity images
âś• Cropping
âś• Reframing
âś• Local self-sufficiency
âś• Coordinate-independent detectionThe final Stage 1 system was frozen as the global multi-carrier baseline. Stage 2 would keep that system available as the control rather than continuing to optimise it indefinitely.
Checkpoint record
The detailed development record contains 43 Stage 1 checkpoints. The list below preserves the research trail without requiring a separate page for every experiment.
View all Stage 1 checkpoints
| Checkpoint | Focus |
|---|---|
| CP001 | Build the image pipeline |
| CP002 | Add and detect a fixed pixel signal |
| CP003 | Test JPEG robustness |
| CP004 | Test resize robustness |
| CP005 | Build a resize-resistant signal |
| CP006 | Test many real images |
| CP007 | Honest holdout evaluation |
| CP008 | Export the failure cases |
| CP009 | Build an adaptive signal |
| CP010 | Benchmark the adaptive signal |
| CP011 | Adaptive encoder with a simple decoder |
| CP012 | Analyse signal capacity versus detection |
| CP013 | Capacity-normalised encoding |
| CP014 | Build the first learned decoder |
| CP015 | Pattern-aware learned decoder |
| CP016 | Test generalisation to unseen images |
| CP017 | Capacity versus learned detection |
| CP018 | Capacity-aware decoder |
| CP019 | Identify which images were rescued |
| CP020 | Decoder ensemble evaluation |
| CP021 | Inspect learned-decoder failures |
| CP022 | Hard-negative training |
| CP023 | Select checkpoints by the real operating metric |
| CP024 | Fresh selection, calibration and evaluation sets |
| CP025 | Fresh transformation robustness |
| CP026 | Pairwise signal-learning loss |
| CP027 | Pairwise-model transformation robustness |
| CP028 | Multi-carrier signal comparison |
| CP029 | Train the multi-carrier learned decoder |
| CP030 | Multi-carrier transformation robustness |
| CP031 | Multi-carrier capacity analysis |
| CP032 | Encoder support policy under transformation |
| CP033 | Capacity-weighted pairwise training |
| CP034 | Paired capacity-weighting comparison |
| CP035 | Tri-state decoder calibration |
| CP036 | Tri-state robustness under transformation |
| CP037 | Final reserve evaluation |
| CP038 | Frozen cross-dataset evaluation |
| CP039 | Oxford capacity profile |
| CP040 | Frozen Pascal VOC evaluation |
| CP041 | Crop and spatial-alignment robustness |
| CP042 | Oracle carrier-alignment recovery |
| CP043 | Calibrated oracle-alignment diagnostic |
Next steps
Stage 2 begins from the frozen Stage 1 system and asks whether useful signal evidence still exists inside surviving image regions and whether those pieces of evidence can be found and combined without knowing the original image geometry.