The question to address for this stage:
Can useful protection evidence survive locally when an image is cropped or reframed, and can a blind detector find and combine that evidence without losing false-positive control?
Stage 1 established that the GAPP signal could survive JPEG compression and
frame-preserving resizing, but cropping exposed a deeper problem. The signal and decoder
were still too dependent on the complete image coordinate frame. Stage 2 therefore focused
primarily on the detection side of the system. Rather than immediately redesigning
the signal or encoder, we kept the Stage 1 representation largely frozen and asked whether
useful evidence already survived inside smaller image regions. By the end of Stage 2,
the answer was yes. We had built a blind local-evidence detector that searched across
possible geometries, normalised those responses against ordinary-image behaviour,
learned bounded corrections to the structured evidence, and combined multiple
independently trained models. We also developed the first research architecture
where PROTECTED and NOT_PROTECTED were treated as separate decisions rather than
opposite ends of one classifier. However, the stage did not solve the deeper
representation problem. The signal itself was still globally defined, and the lower
tri-state threshold did not transfer reliably enough across every tested condition.
Starting point
Stage 2 started with the frozen Stage 1 system:
Content-adaptive encoder
+
Three low-frequency RGB carriers
+
Blind capacity-aware decoderThat system had demonstrated strong performance on complete images and frame-preserving transformations. The unresolved problem was cropping. Stage 1 had shown that knowing the crop geometry could recover some lost evidence, but even exact alignment did not restore the original clean-image performance. The learned decoder had become specialised around the original image geometry. This suggested two possibilities:
Possibility A - Cropping physically destroys most of the signal
Possibility B - The signal survives locally, but the detector does not know where or how to find itBefore changing the encoder, Stage 2 needed to determine which explanation was closer to reality. We therefore deliberately kept the Stage 1 signal and encoder stable during most of the stage. Changing the signal, encoder, evidence extractor and detector simultaneously would have made it difficult to determine why a result improved or failed.
What we explored
1. Whether cropped regions still contained usable signal evidence
The first Stage 2 experiments bypassed the Stage 1 learned decoder and examined carrier evidence directly inside local windows. When the correct surviving geometry was known, the result was encouraging. Using aligned local evidence on Pascal VOC:
| Condition | False positives | Protected detection | Pair ordering |
|---|---|---|---|
| Clean | 0.98% | 62.70% | 100% |
| Centre crop 90% | 0.78% | 61.13% | 100% |
| Top-left crop 90% | 0.98% | 58.01% | 100% |
| Bottom-right crop 90% | 0.59% | 58.59% | 100% |
| Centre crop 75% | 0.39% | 46.09% | 100% |
The mean protected score lift also remained relatively stable across the crop conditions. This established that cropping did not necessarily erase the GAPP signal. Useful evidence still existed inside surviving image regions. The problem was increasingly becoming one of blind localisation and interpretation.
2. Blind geometric search introduced a multiple-hypothesis problem
A real detector cannot know which crop occurred. We therefore introduced a
bank of 19 geometric hypotheses covering different positions, scales and crop
configurations. The detector searched them blindly. This worked geometrically: the
correct or approximately correct hypothesis was often among the strongest candidates. But
a new statistical problem appeared.
More hypotheses searched
β
More opportunities for ordinary image content
to accidentally resemble the signal
β
Higher ordinary-image score tail
β
Higher threshold required
β
Lower protected detectionEven clean images suffered because searching many hypotheses gave ordinary content more opportunities to produce an accidental strong match. The problem was therefore not simply to find the highest correlation. A local detector also needed to understand how unusual a correlation was relative to ordinary images under that same hypothesis.
3. We normalised each hypothesis against ordinary-image behaviour
Stage 2 introduced a shared null model. Rather than comparing raw local correlation values directly, each hypothesis was interpreted relative to the score distribution produced by ordinary images. The result was approximately:
raw local response
β
ordinary-image null normalisation
β
normalised evidence
β
aggregate several useful regionsThis substantially reduced the advantage of naturally high-scoring hypotheses. By Checkpoint 47, the first genuinely blind local-evidence baseline used:
19 geometric hypotheses
+
local spatial windows
+
shared ordinary-image null model
+
one shared thresholdOn Pascal VOC, it reached:
Protected detection: 40.00%
False positives: 0.70%
Pair ordering: 91.88%The 91.88% pair ordering was especially informative. Protected copies usually received more evidence
than their matching ordinary copies, even when many of them did not cross the strict decision threshold.
The signal was generally pushing scores in the correct direction. The remaining problem was separating
the overlapping score tails.
4. We tested whether learning could improve evidence aggregation
The next branch asked whether a learned model could interpret the full local-evidence tensor better than handcrafted aggregation. The first attempt failed. A learned local accumulator was allowed to replace the existing structured score entirely. Training collapsed even though the evidence tensor appeared informative. A small overfitting experiment then showed that the architecture could memorise a tiny dataset. The failure was therefore not simply a lack of information. The more important architectural lesson was that we were discarding an already useful baseline. The next model changed the relationship:
Handcrafted evidence score + bounded learned correction = final scoreInstead of asking a neural model to rediscover the complete detector, the structured carrier-aware baseline remained intact. The model only learned when that baseline should be adjusted. This became the residual evidence accumulator.
5. We explored the scale of local evidence
The early local detector used 64 Γ 64 windows. A multiscale diagnostic later showed that smaller
windows contained stronger useful evidence. The final selected representation used:
Window size: 32 Γ 32
Stride: 16At 256 Γ 256 evaluation resolution, this produced 225 overlapping local regions for each geometric hypothesis.
This was computationally much heavier, but it provided a stronger local-evidence representation.
The resulting CP058 residual detector became the main Stage 2 single-model reference.
6. We investigated why false positives appeared under new domains
As the detector improved, a recurring problem became more visible. Some learned residual corrections behaved differently when transferred from Caltech to Flowers. A false positive could arise from:
moderately high ordinary-image evidence + large positive learned correction = false PROTECTED decisionSeveral approaches attempted to detect these mistakes after the final score had already been produced. We tested:
- baseline support gates;
- hypothesis-consensus rules;
- prominence thresholds;
- borderline abstention;
- joint-confidence rules;
- a learned confidence head;
- hard-negative confidence training.
These diagnostics identified real structure, but the trade-offs were poor. Preventing one false positive often required abstaining on several correctly protected images. The more successful intervention was therefore moved back into training. Instead of trying to recognise bad positive corrections afterward, the model was penalised when it applied large positive corrections to ordinary images that already had unusually high baseline evidence. That produced the tail-weighted training objective used later in the stage.
7. We tested whether the training result depended on one lucky seed
A single good model is not strong evidence that the training procedure itself is reliable. The selected tail-weighted configuration was therefore retrained across five random seeds. All five satisfied the predefined internal retention requirements. Across the five seeds:
Validation protected detection:
60.36% Β± 0.36%
Pairwise ordering:
98.50% Β± 0.26%This suggested that the training recipe itself was stable. However, the individual models still differed enough in their external false-positive behaviour that selecting one seed remained undesirable. That motivated ensembling.
8. We compared different ensemble aggregation rules
Five independently trained residual models were combined. Mean correction improved recall, but one external condition exceeded the preferred false-positive envelope. The aggregation rule was therefore selected using validation data rather than whichever rule happened to perform best externally. Median correction was chosen. The median reduced the influence of unusually aggressive ensemble members while retaining the improvement shared across the group. This became CP082.
9. We separated the two tri-state decisions
Stage 1 had already introduced the idea of:
NOT_PROTECTED
INDETERMINATE
PROTECTEDStage 2 revealed that these should not necessarily be produced from one score using two thresholds.
CP082 was better at identifying PROTECTED images, but its lower score tail was extremely conservative.
On the full Flowers test set, using CP082 alone for both sides resulted in:
Ordinary NOT_PROTECTED: 1.22%
Ordinary INDETERMINATE: 97.49%The earlier CP058 score transferred better when identifying ordinary images. We therefore treated the two decisions as separate statistical questions.
NOT_PROTECTED uses CP058 lower-tail evidence
PROTECTED uses CP082 median-ensemble evidence
Anything else becomes INDETERMINATEThis became the CP085 hybrid tri-state architecture.
What we built
The final Stage 2 research system retained the Stage 1 signal and encoder but substantially changed the detector. A simplified view is:
Received image
β
19 geometry hypotheses
β
Overlapping 32 Γ 32 local windows
stride 16
β
Carrier-aware local evidence
β
Shared null normalisation
β
Structured baseline score
β
Bounded learned residual correction
β
Five independently trained models
β
Median correction aggregation
β
CP082 PROTECTED scoreThe frozen binary detector was:
| Component | Final Stage 2 configuration |
|---|---|
| Signal | Frozen Stage 1 multi-carrier signal |
| Encoder | Frozen Stage 1 content-adaptive encoder |
| Geometry search | 19 hypotheses |
| Local window | 32 Γ 32 |
| Window stride | 16 |
| Local evidence | Null-normalised carrier evidence |
| Learned component | Bounded residual correction |
| Correction cap | 3.0 |
| Training | Tail-weighted ordinary correction penalty |
| Training seeds | 5 |
| Ensemble | Median residual correction |
| Binary detector | CP082 |
The Stage 2 hybrid tri-state architecture used two different score sources:
CP058 score β€ 0.027639 = NOT_PROTECTED
CP082 score β₯ 3.348145 = PROTECTED
Otherwise = INDETERMINATE
If both branches somehow contradict = INDETERMINATEThe two thresholds remained experimental calibration values rather than permanent protocol constants. The important architectural change was the separation itself:
Failure to establish PROTECTED was no longer treated as evidence of NOT_PROTECTED.
Key results
Blind local evidence baseline
The first deployment-shaped local detector reached:
Pascal VOC diagnostic sample
Protected detection: 40.00%
False positives: 0.70%
Pair ordering: 91.88%This established that blind local evidence extraction was viable, although natural-image tails still limited thresholded recall.
Full Flowers102 evaluation
The final comparison between the single-model CP058 reference and the CP082 median ensemble
used the complete 6,149-image Flowers102 test split across nine processing conditions.
| Detector | False positives | Protected detection | Maximum FPR | Pair ordering |
|---|---|---|---|---|
| CP058 | 1.30% | 52.66% | 1.53% | 98.57% |
| CP082 median | 1.29% | 53.73% | 1.61% | 99.03% |
The numerical improvement was small, but it reproduced across the full test set. A paired image-level bootstrap estimated:
CP082 β CP058 protected-detection improvement
95% interval:
+0.74 to +1.40 percentage pointsThe false-positive difference was statistically indistinguishable from zero. The appropriate interpretation was therefore:
CP082 produced a small but reproducible protected-detection improvement while maintaining approximately the same pooled false-positive rate.
Full Flowers hybrid tri-state result
The CP085 hybrid architecture used CP058 for NOT_PROTECTED and CP082 for PROTECTED. Across the complete
Flowers test split:
For ordinary images,
NOT_PROTECTED 14.48%
INDETERMINATE 84.23%
PROTECTED 1.29%
---
For protected images,
NOT_PROTECTED 0.63%
INDETERMINATE 45.64%
PROTECTED 53.73%
---
Overall:
Definitive coverage: 35.06%
Definitive accuracy: 97.26%
Wrong definitive rate: 0.96%
Conflict rate: 0.00%Compared with CP058 alone, the hybrid slightly improved protected detection and coverage without increasing the overall wrong-definitive rate.
Uncontrolled Wikimedia Commons benchmark
The frozen system was also evaluated on 500 random Wikimedia Commons images. This dataset was
intentionally less curated than the earlier academic benchmarks and contained:
450 JPEG
22 PNG
19 TIFF
8 SVG
1 GIFThe images included photographs, scans, illustrations, diagrams and unusually shaped web media. Across nine processing conditions:
| Detector | False positives | Protected detection | Maximum FPR | Minimum TPR | Pair ordering |
|---|---|---|---|---|---|
| CP058 | 0.22% | 66.93% | 0.40% | 57.40% | 99.20% |
| CP082 median | 0.20% | 71.82% | 0.40% | 61.80% | 99.22% |
CP082 improved protected detection in every tested condition while maintaining very low false-PROTECTED rates. The hybrid tri-state system produced:
For ordinary images,
NOT_PROTECTED 18.33%
INDETERMINATE 81.47%
PROTECTED 0.20%
---
For protected images,
NOT_PROTECTED 0.91%
INDETERMINATE 27.27%
PROTECTED 71.82%
---
Overall:
Definitive coverage: 45.63%
Definitive accuracy: 98.78%
Wrong definitive rate: 0.56%
Conflict rate: 0.00%The main qualification appeared in the lower decision. The pooled false-NOT_PROTECTED rate was 0.91% but
the worst individual condition reached 1.60%, against the intended 1% condition-level target. That
prevented the lower threshold from being fully frozen. The Wikimedia result was encouraging, but it should
not be treated as a complete real-world internet benchmark. It did not reproduce every social platform,
screenshot path, colour-management chain, device pipeline or generative transformation.
What worked
Several Stage 2 approaches produced results strong enough to retain.
-
Local carrier evidence survived cropping: When the surviving geometry was known, cropped regions still contained strong signal evidence. Cropping was therefore not simply erasing everything.
-
Null normalisation made blind search usable: Interpreting each hypothesis relative to ordinary-image behaviour reduced the statistical advantage of naturally high-scoring locations.
-
Smaller overlapping windows improved the evidence representation: The
32 Γ 32, stride-16representation outperformed the earlier64 Γ 64local baseline. -
Preserving the structured baseline worked better than replacing it: The residual accumulator learned bounded corrections to an explicit carrier-aware score rather than trying to rediscover the entire detector.
-
Training against the difficult ordinary-image tail improved transfer: Penalising positive corrections specifically on high-baseline ordinary images was more effective than adding increasingly complicated post-hoc confidence rules.
-
The tail-weighted training recipe reproduced across seeds: Five independent runs produced tightly grouped validation behaviour, reducing concern that the result depended on one fortunate initialization.
-
Median ensembling reduced seed-specific instability: Combining five residual models using the median produced a more stable external operating point than relying on one selected model or a simple mean.
-
Separate positive and negative decision branches were useful: CP058 and CP082 behaved differently in the lower and upper score tails. Combining their strengths produced a better tri-state architecture than forcing one score to make both decisions.
-
Frozen detectors transferred to uncontrolled web content: The Wikimedia benchmark produced higher protected detection and lower false positives than Flowers, although that result remains specific to the tested population.
What did not work
Several branches were informative but were not promoted into the Stage 2 system.
- Learned local evidence replacement: The first learned local accumulator attempted to replace the handcrafted score entirely and collapsed during full training.
Lesson: useful structured evidence should not be discarded simply because learning is introduced.
- Raw maximum over geometry hypotheses: Searching many hypotheses and selecting the strongest raw response inflated ordinary-image tails.
Lesson: blind search has a statistical cost that must be calibrated explicitly.
- Larger
64 Γ 64windows as the final representation: They contained useful evidence but were weaker than the later32 Γ 32representation.
Lesson: smaller overlapping regions provided a better local-evidence basis.
- Simple platform-wide threshold tightening: Raising one global threshold improved internal false-positive control but did not fully solve external domain drift.
Lesson: threshold adjustment alone cannot repair distribution-dependent learned corrections.
- Positive-correction support gates: Requiring a minimum baseline score before allowing positive correction did not cleanly distinguish false positives from protected images that genuinely benefited from correction.
Lesson: a one-dimensional support floor was too crude.
- Hypothesis prominence and consensus rules: False positives did not consistently look like isolated spikes, and true detections were not always broadly distributed.
Lesson: the spatial and hypothesis evidence structure was more complex than a simple consensus rule.
- Borderline abstention and joint-confidence rules: These reduced false positives but required abstaining on too many correct protected decisions.
Lesson: post-hoc confidence filters had an inefficient recall cost.
- Learned confidence head: The model achieved strong overall AUC but performed poorly in the narrow near-threshold region that actually mattered operationally.
Lesson: good global ranking does not guarantee useful tail separation.
- Hard-negative confidence training: Adding more near-threshold ordinary examples did not make the confidence layer sufficiently efficient.
Lesson: the failure could not be solved simply by expanding the confidence-head negative pool.
- Global asymmetric correction penalty alone: Penalising ordinary positive corrections was promising, but its learned correction distribution still drifted across domains.
Lesson: the difficult high-baseline tail needed explicit treatment.
- Mean ensemble aggregation: Mean correction improved recall but allowed one condition to exceed the intended external false-positive envelope.
Lesson: robust aggregation mattered because individual models were not equally conservative.
- Single-score tri-state decoding: CP082 worked well on the upper tail but classified almost every ordinary image as
INDETERMINATEwhen used for the lower decision.
Lesson:
PROTECTEDandNOT_PROTECTEDshould be treated as different statistical problems.
The signal itself remained globally defined
Despite the detector improvements, the deeper Stage 1 architectural limitation was not removed. Stage 2 could search for local evidence much more effectively, but that evidence still came from the original globally defined carrier system. The system therefore worked roughly like:
Global signal
β
crop or reframe
β
surviving fragments of global signal
β
search many possible geometries
β
recover local evidenceThis is different from an encoder where each sufficiently large region is intentionally designed to contain self-sufficient protection evidence. Stage 2 improved the recovery of local evidence. It did not yet make the signal itself intrinsically local.
What Stage 2 taught us
-
The first major lesson was that the Stage 1 crop failure had been partly mischaracterised. The signal was not simply disappearing. Instead useful protection evidence often survived inside cropped regions, but finding it blindly introduced a difficult statistical search problem**. If cropping had physically destroyed the carrier, the obvious next move would have been to increase robustness or signal strength. Instead, Stage 2 showed that a large part of the problem involved representation, search and evidence aggregation.
-
The second major lesson was that structured and learned methods were complementary. Pure handcrafted correlation had useful inductive structure but insufficient separation. Pure learned replacement discarded information we already knew was meaningful. The strongest Stage 2 detector combined both:
explicit carrier evidence
+
ordinary-image calibration
+
learned bounded correction-
The third lesson was that false-positive behaviour must be studied in the score tails rather than through average accuracy alone. Several models showed excellent ordering while still producing unacceptable errors around the decision threshold. Training distribution, domain shift and the behaviour of rare difficult ordinary images mattered as much as aggregate model quality.
-
The fourth lesson was that
PROTECTEDandNOT_PROTECTEDwere not naturally symmetric. A detector that is strong at establishing the presence of protection evidence may still be poor at establishing its absence. That led to a more conservative design principle:
No convincing protection evidence
β
proof of no protectionThe INDETERMINATE state therefore became increasingly important rather than something to minimise at any cost.
- Finally, Stage 2 showed that better local decoding alone was unlikely to finish the problem. The detector had become considerably more capable, but the underlying signal still inherited the global geometry of Stage 1. The next architectural question was therefore: Can the signal itself be redesigned so that surviving local regions carry independently useful evidence from the beginning?
Outcome
STAGE 2
COMPLETE_WITH_PROVISIONAL_TRISTATE_THRESHOLD
Established:
β Cropped regions retain useful local carrier evidence
β Blind local evidence extraction is possible
β Shared null normalisation works across geometry hypotheses
β 32 Γ 32 overlapping evidence improves the local representation
β Baseline-preserving residual learning works
β Tail-weighted training is stable across five seeds
β Median ensembling improves detector stability
β Frozen binary detector transfers to Flowers and Wikimedia
β Separate PROTECTED and NOT_PROTECTED branches are useful
β Hybrid tri-state architecture produces no observed branch conflicts
Frozen:
β CP058 single-model lower branch
β CP078 training recipe
β CP082 five-model median PROTECTED detector
β CP085 hybrid tri-state architecture
β Wikimedia 500-image benchmark manifest
Provisional:
β³ NOT_PROTECTED threshold policy
Unresolved:
β³ High INDETERMINATE rate
β³ Condition-level lower-threshold transfer
β³ Five-model inference complexity
β³ Full real-world platform coverage
β³ Spatial evidence interactions
β Intrinsically local signal representation
β Locally self-sufficient encoding
β Full crop/reframe independence
β Deliberate removal robustness
β Generative transformation robustnessThe Stage 2 binary detector was frozen as CP082, the five-model median residual-correction ensemble. The hybrid CP085 structure was also retained as the Stage 2 research architecture:
CP058 lower branch
+
CP082 upper branch
+
INDETERMINATE between themHowever, the lower threshold remained provisional because one Wikimedia condition reached
a 1.60% false-NOT_PROTECTED rate against the intended 1% condition-level limit. No further
optimisation should be performed against the already examined Flowers or Wikimedia populations.
Checkpoint record
The detailed development record contains 45 Stage 2 checkpoints. The list below preserves the research trail without requiring a separate page for every experiment.
View all Stage 2 checkpoints
| Checkpoint | Focus |
|---|---|
| CP044 | Local evidence baseline and spatial evidence mapping |
| CP045 | Blind local alignment hypothesis bank |
| CP046 | Null-normalised hypothesis evidence accumulation |
| CP047 | Shared null model and one blind threshold |
| CP048 | Learned local evidence accumulator |
| CP049 | Tiny-set overfit test |
| CP050 | Baseline-preserving residual evidence accumulator |
| CP051 | Frozen external evaluation and correction audit |
| CP052 | Residual correction-cap ablation |
| CP053 | Frozen full-VOC comparison |
| CP054 | Extended correction-cap boundary |
| CP055 | Multi-seed residual-cap stability |
| CP056 | Multiscale local-evidence diagnostic |
| CP057 | Frozen 32-pixel local baseline transfer |
| CP058 | Window-32 residual accumulator |
| CP059 | Window-32 residual multi-seed stability |
| CP060 | Frozen external window-32 evaluation |
| CP061 | Frozen full Oxford Flowers confirmation |
| CP062 | Frozen platform-transform stress test |
| CP063 | Platform threshold-drift decomposition |
| CP064 | Platform-aware shared threshold calibration |
| CP065 | External transfer of the platform-aware threshold |
| CP066 | Cross-domain false-positive decomposition |
| CP067 | Positive-correction support-gate diagnostic |
| CP068 | Hypothesis-consensus false-positive audit |
| CP069 | Operating-point transition audit |
| CP070 | Borderline low-prominence abstention diagnostic |
| CP071 | Joint-confidence abstention diagnostic |
| CP072 | Learned positive-decision confidence head |
| CP073 | Hard-negative confidence training |
| CP074 | Asymmetric ordinary correction penalty |
| CP075 | Matched-FPR and calibration-drift diagnostic |
| CP076 | Score-scale and false-positive mechanism audit |
| CP077 | Conditional correction drift audit |
| CP078 | High-baseline-tail correction penalty |
| CP079 | Tail-weighted multiseed stability |
| CP080 | Tail-weighted multiseed external transfer |
| CP081 | Five-seed mean-correction ensemble |
| CP082 | Robust correction aggregation |
| CP083 | Full Flowers fixed-threshold confirmation |
| CP084 | Tri-state calibration |
| CP085 | Hybrid tri-state decoder |
| CP086 | End-to-end evidence-density maps |
| CP087 | Hostile internet benchmark |
| CP088 | Formal Stage 2 review and artifact freeze |
Next steps
Stage 3 should begin from the frozen Stage 2 detector rather than continuing to extract marginal improvements from the same global carrier. The central research question becomes whether GAPP can move from recovering fragments of a globally defined signal to an architecture where local regions are intentionally designed to carry useful protection evidence independently. The next stage should therefore investigate:
- locally self-sufficient signal representations;
- overlapping and redundant local encoding;
- stronger crop and reframing tolerance;
- whether the existing detector can remain useful as a frozen baseline;
- perceptual quality of the new representation;
- real screenshot and platform-processing channels;
- generative preprocessing and reconstruction;
- deliberate signal removal and spoofing;
- the computational cost of blind search;
- more rigorous tri-state decision semantics.
The Stage 2 architecture should remain available as the reference system while those questions are investigated.