πŸ› οΈ Current status: Stage 4 | Viability boundaries and protocol hardening Β· Updated 24 August 2026
Stage 2: Local evidence & detection

The question to address for this stage:

Can useful protection evidence survive locally when an image is cropped or reframed, and can a blind detector find and combine that evidence without losing false-positive control?

Stage 1 established that the GAPP signal could survive JPEG compression and frame-preserving resizing, but cropping exposed a deeper problem. The signal and decoder were still too dependent on the complete image coordinate frame. Stage 2 therefore focused primarily on the detection side of the system. Rather than immediately redesigning the signal or encoder, we kept the Stage 1 representation largely frozen and asked whether useful evidence already survived inside smaller image regions. By the end of Stage 2, the answer was yes. We had built a blind local-evidence detector that searched across possible geometries, normalised those responses against ordinary-image behaviour, learned bounded corrections to the structured evidence, and combined multiple independently trained models. We also developed the first research architecture where PROTECTED and NOT_PROTECTED were treated as separate decisions rather than opposite ends of one classifier. However, the stage did not solve the deeper representation problem. The signal itself was still globally defined, and the lower tri-state threshold did not transfer reliably enough across every tested condition.


Starting point

Stage 2 started with the frozen Stage 1 system:

Content-adaptive encoder
        +
Three low-frequency RGB carriers
        +
Blind capacity-aware decoder

That system had demonstrated strong performance on complete images and frame-preserving transformations. The unresolved problem was cropping. Stage 1 had shown that knowing the crop geometry could recover some lost evidence, but even exact alignment did not restore the original clean-image performance. The learned decoder had become specialised around the original image geometry. This suggested two possibilities:

Possibility A - Cropping physically destroys most of the signal

Possibility B - The signal survives locally, but the detector does not know where or how to find it

Before changing the encoder, Stage 2 needed to determine which explanation was closer to reality. We therefore deliberately kept the Stage 1 signal and encoder stable during most of the stage. Changing the signal, encoder, evidence extractor and detector simultaneously would have made it difficult to determine why a result improved or failed.


What we explored

1. Whether cropped regions still contained usable signal evidence

The first Stage 2 experiments bypassed the Stage 1 learned decoder and examined carrier evidence directly inside local windows. When the correct surviving geometry was known, the result was encouraging. Using aligned local evidence on Pascal VOC:

ConditionFalse positivesProtected detectionPair ordering
Clean0.98%62.70%100%
Centre crop 90%0.78%61.13%100%
Top-left crop 90%0.98%58.01%100%
Bottom-right crop 90%0.59%58.59%100%
Centre crop 75%0.39%46.09%100%

The mean protected score lift also remained relatively stable across the crop conditions. This established that cropping did not necessarily erase the GAPP signal. Useful evidence still existed inside surviving image regions. The problem was increasingly becoming one of blind localisation and interpretation.

2. Blind geometric search introduced a multiple-hypothesis problem

A real detector cannot know which crop occurred. We therefore introduced a bank of 19 geometric hypotheses covering different positions, scales and crop configurations. The detector searched them blindly. This worked geometrically: the correct or approximately correct hypothesis was often among the strongest candidates. But a new statistical problem appeared.

More hypotheses searched
        ↓
More opportunities for ordinary image content
to accidentally resemble the signal
        ↓
Higher ordinary-image score tail
        ↓
Higher threshold required
        ↓
Lower protected detection

Even clean images suffered because searching many hypotheses gave ordinary content more opportunities to produce an accidental strong match. The problem was therefore not simply to find the highest correlation. A local detector also needed to understand how unusual a correlation was relative to ordinary images under that same hypothesis.

3. We normalised each hypothesis against ordinary-image behaviour

Stage 2 introduced a shared null model. Rather than comparing raw local correlation values directly, each hypothesis was interpreted relative to the score distribution produced by ordinary images. The result was approximately:

raw local response
        ↓
ordinary-image null normalisation
        ↓
normalised evidence
        ↓
aggregate several useful regions

This substantially reduced the advantage of naturally high-scoring hypotheses. By Checkpoint 47, the first genuinely blind local-evidence baseline used:

19 geometric hypotheses 
+ 
local spatial windows 
+ 
shared ordinary-image null model 
+ 
one shared threshold

On Pascal VOC, it reached:

Protected detection:  40.00%
False positives:        0.70%
Pair ordering:         91.88%

The 91.88% pair ordering was especially informative. Protected copies usually received more evidence than their matching ordinary copies, even when many of them did not cross the strict decision threshold. The signal was generally pushing scores in the correct direction. The remaining problem was separating the overlapping score tails.

4. We tested whether learning could improve evidence aggregation

The next branch asked whether a learned model could interpret the full local-evidence tensor better than handcrafted aggregation. The first attempt failed. A learned local accumulator was allowed to replace the existing structured score entirely. Training collapsed even though the evidence tensor appeared informative. A small overfitting experiment then showed that the architecture could memorise a tiny dataset. The failure was therefore not simply a lack of information. The more important architectural lesson was that we were discarding an already useful baseline. The next model changed the relationship:

Handcrafted evidence score + bounded learned correction = final score

Instead of asking a neural model to rediscover the complete detector, the structured carrier-aware baseline remained intact. The model only learned when that baseline should be adjusted. This became the residual evidence accumulator.

5. We explored the scale of local evidence

The early local detector used 64 Γ— 64 windows. A multiscale diagnostic later showed that smaller windows contained stronger useful evidence. The final selected representation used:

Window size:  32 Γ— 32
Stride:       16

At 256 Γ— 256 evaluation resolution, this produced 225 overlapping local regions for each geometric hypothesis. This was computationally much heavier, but it provided a stronger local-evidence representation. The resulting CP058 residual detector became the main Stage 2 single-model reference.

6. We investigated why false positives appeared under new domains

As the detector improved, a recurring problem became more visible. Some learned residual corrections behaved differently when transferred from Caltech to Flowers. A false positive could arise from:

moderately high ordinary-image evidence + large positive learned correction = false PROTECTED decision

Several approaches attempted to detect these mistakes after the final score had already been produced. We tested:

  • baseline support gates;
  • hypothesis-consensus rules;
  • prominence thresholds;
  • borderline abstention;
  • joint-confidence rules;
  • a learned confidence head;
  • hard-negative confidence training.

These diagnostics identified real structure, but the trade-offs were poor. Preventing one false positive often required abstaining on several correctly protected images. The more successful intervention was therefore moved back into training. Instead of trying to recognise bad positive corrections afterward, the model was penalised when it applied large positive corrections to ordinary images that already had unusually high baseline evidence. That produced the tail-weighted training objective used later in the stage.

7. We tested whether the training result depended on one lucky seed

A single good model is not strong evidence that the training procedure itself is reliable. The selected tail-weighted configuration was therefore retrained across five random seeds. All five satisfied the predefined internal retention requirements. Across the five seeds:

Validation protected detection:
60.36% Β± 0.36%

Pairwise ordering:
98.50% Β± 0.26%

This suggested that the training recipe itself was stable. However, the individual models still differed enough in their external false-positive behaviour that selecting one seed remained undesirable. That motivated ensembling.

8. We compared different ensemble aggregation rules

Five independently trained residual models were combined. Mean correction improved recall, but one external condition exceeded the preferred false-positive envelope. The aggregation rule was therefore selected using validation data rather than whichever rule happened to perform best externally. Median correction was chosen. The median reduced the influence of unusually aggressive ensemble members while retaining the improvement shared across the group. This became CP082.

9. We separated the two tri-state decisions

Stage 1 had already introduced the idea of:

NOT_PROTECTED
INDETERMINATE
PROTECTED

Stage 2 revealed that these should not necessarily be produced from one score using two thresholds. CP082 was better at identifying PROTECTED images, but its lower score tail was extremely conservative. On the full Flowers test set, using CP082 alone for both sides resulted in:

Ordinary NOT_PROTECTED:  1.22%
Ordinary INDETERMINATE: 97.49%

The earlier CP058 score transferred better when identifying ordinary images. We therefore treated the two decisions as separate statistical questions.

NOT_PROTECTED uses CP058 lower-tail evidence

PROTECTED uses CP082 median-ensemble evidence

Anything else becomes INDETERMINATE

This became the CP085 hybrid tri-state architecture.


What we built

The final Stage 2 research system retained the Stage 1 signal and encoder but substantially changed the detector. A simplified view is:

Received image
      ↓
19 geometry hypotheses
      ↓
Overlapping 32 Γ— 32 local windows
stride 16
      ↓
Carrier-aware local evidence
      ↓
Shared null normalisation
      ↓
Structured baseline score
      ↓
Bounded learned residual correction
      ↓
Five independently trained models
      ↓
Median correction aggregation
      ↓
CP082 PROTECTED score

The frozen binary detector was:

ComponentFinal Stage 2 configuration
SignalFrozen Stage 1 multi-carrier signal
EncoderFrozen Stage 1 content-adaptive encoder
Geometry search19 hypotheses
Local window32 Γ— 32
Window stride16
Local evidenceNull-normalised carrier evidence
Learned componentBounded residual correction
Correction cap3.0
TrainingTail-weighted ordinary correction penalty
Training seeds5
EnsembleMedian residual correction
Binary detectorCP082

The Stage 2 hybrid tri-state architecture used two different score sources:

CP058 score ≀ 0.027639 = NOT_PROTECTED

CP082 score β‰₯ 3.348145 = PROTECTED

Otherwise = INDETERMINATE

If both branches somehow contradict = INDETERMINATE

The two thresholds remained experimental calibration values rather than permanent protocol constants. The important architectural change was the separation itself:

Failure to establish PROTECTED was no longer treated as evidence of NOT_PROTECTED.


Key results

Blind local evidence baseline

The first deployment-shaped local detector reached:

Pascal VOC diagnostic sample

Protected detection:  40.00%
False positives:        0.70%
Pair ordering:         91.88%

This established that blind local evidence extraction was viable, although natural-image tails still limited thresholded recall.

Full Flowers102 evaluation

The final comparison between the single-model CP058 reference and the CP082 median ensemble used the complete 6,149-image Flowers102 test split across nine processing conditions.

DetectorFalse positivesProtected detectionMaximum FPRPair ordering
CP0581.30%52.66%1.53%98.57%
CP082 median1.29%53.73%1.61%99.03%

The numerical improvement was small, but it reproduced across the full test set. A paired image-level bootstrap estimated:

CP082 βˆ’ CP058 protected-detection improvement
95% interval:
+0.74 to +1.40 percentage points

The false-positive difference was statistically indistinguishable from zero. The appropriate interpretation was therefore:

CP082 produced a small but reproducible protected-detection improvement while maintaining approximately the same pooled false-positive rate.

Full Flowers hybrid tri-state result

The CP085 hybrid architecture used CP058 for NOT_PROTECTED and CP082 for PROTECTED. Across the complete Flowers test split:

For ordinary images,
NOT_PROTECTED    14.48%
INDETERMINATE    84.23%
PROTECTED         1.29%
---
For protected images,
NOT_PROTECTED     0.63%
INDETERMINATE    45.64%
PROTECTED        53.73%
---
Overall:
Definitive coverage:     35.06%
Definitive accuracy:     97.26%
Wrong definitive rate:    0.96%
Conflict rate:             0.00%

Compared with CP058 alone, the hybrid slightly improved protected detection and coverage without increasing the overall wrong-definitive rate.

Uncontrolled Wikimedia Commons benchmark

The frozen system was also evaluated on 500 random Wikimedia Commons images. This dataset was intentionally less curated than the earlier academic benchmarks and contained:

450 JPEG
22 PNG
19 TIFF
8 SVG
1 GIF

The images included photographs, scans, illustrations, diagrams and unusually shaped web media. Across nine processing conditions:

DetectorFalse positivesProtected detectionMaximum FPRMinimum TPRPair ordering
CP0580.22%66.93%0.40%57.40%99.20%
CP082 median0.20%71.82%0.40%61.80%99.22%

CP082 improved protected detection in every tested condition while maintaining very low false-PROTECTED rates. The hybrid tri-state system produced:

For ordinary images,
NOT_PROTECTED    18.33%
INDETERMINATE    81.47%
PROTECTED         0.20%
---
For protected images,
NOT_PROTECTED     0.91%
INDETERMINATE    27.27%
PROTECTED        71.82%
---
Overall:
Definitive coverage:     45.63%
Definitive accuracy:     98.78%
Wrong definitive rate:    0.56%
Conflict rate:             0.00%

The main qualification appeared in the lower decision. The pooled false-NOT_PROTECTED rate was 0.91% but the worst individual condition reached 1.60%, against the intended 1% condition-level target. That prevented the lower threshold from being fully frozen. The Wikimedia result was encouraging, but it should not be treated as a complete real-world internet benchmark. It did not reproduce every social platform, screenshot path, colour-management chain, device pipeline or generative transformation.


What worked

Several Stage 2 approaches produced results strong enough to retain.

  1. Local carrier evidence survived cropping: When the surviving geometry was known, cropped regions still contained strong signal evidence. Cropping was therefore not simply erasing everything.

  2. Null normalisation made blind search usable: Interpreting each hypothesis relative to ordinary-image behaviour reduced the statistical advantage of naturally high-scoring locations.

  3. Smaller overlapping windows improved the evidence representation: The 32 Γ— 32, stride-16 representation outperformed the earlier 64 Γ— 64 local baseline.

  4. Preserving the structured baseline worked better than replacing it: The residual accumulator learned bounded corrections to an explicit carrier-aware score rather than trying to rediscover the entire detector.

  5. Training against the difficult ordinary-image tail improved transfer: Penalising positive corrections specifically on high-baseline ordinary images was more effective than adding increasingly complicated post-hoc confidence rules.

  6. The tail-weighted training recipe reproduced across seeds: Five independent runs produced tightly grouped validation behaviour, reducing concern that the result depended on one fortunate initialization.

  7. Median ensembling reduced seed-specific instability: Combining five residual models using the median produced a more stable external operating point than relying on one selected model or a simple mean.

  8. Separate positive and negative decision branches were useful: CP058 and CP082 behaved differently in the lower and upper score tails. Combining their strengths produced a better tri-state architecture than forcing one score to make both decisions.

  9. Frozen detectors transferred to uncontrolled web content: The Wikimedia benchmark produced higher protected detection and lower false positives than Flowers, although that result remains specific to the tested population.


What did not work

Several branches were informative but were not promoted into the Stage 2 system.

  1. Learned local evidence replacement: The first learned local accumulator attempted to replace the handcrafted score entirely and collapsed during full training.

Lesson: useful structured evidence should not be discarded simply because learning is introduced.

  1. Raw maximum over geometry hypotheses: Searching many hypotheses and selecting the strongest raw response inflated ordinary-image tails.

Lesson: blind search has a statistical cost that must be calibrated explicitly.

  1. Larger 64 Γ— 64 windows as the final representation: They contained useful evidence but were weaker than the later 32 Γ— 32 representation.

Lesson: smaller overlapping regions provided a better local-evidence basis.

  1. Simple platform-wide threshold tightening: Raising one global threshold improved internal false-positive control but did not fully solve external domain drift.

Lesson: threshold adjustment alone cannot repair distribution-dependent learned corrections.

  1. Positive-correction support gates: Requiring a minimum baseline score before allowing positive correction did not cleanly distinguish false positives from protected images that genuinely benefited from correction.

Lesson: a one-dimensional support floor was too crude.

  1. Hypothesis prominence and consensus rules: False positives did not consistently look like isolated spikes, and true detections were not always broadly distributed.

Lesson: the spatial and hypothesis evidence structure was more complex than a simple consensus rule.

  1. Borderline abstention and joint-confidence rules: These reduced false positives but required abstaining on too many correct protected decisions.

Lesson: post-hoc confidence filters had an inefficient recall cost.

  1. Learned confidence head: The model achieved strong overall AUC but performed poorly in the narrow near-threshold region that actually mattered operationally.

Lesson: good global ranking does not guarantee useful tail separation.

  1. Hard-negative confidence training: Adding more near-threshold ordinary examples did not make the confidence layer sufficiently efficient.

Lesson: the failure could not be solved simply by expanding the confidence-head negative pool.

  1. Global asymmetric correction penalty alone: Penalising ordinary positive corrections was promising, but its learned correction distribution still drifted across domains.

Lesson: the difficult high-baseline tail needed explicit treatment.

  1. Mean ensemble aggregation: Mean correction improved recall but allowed one condition to exceed the intended external false-positive envelope.

Lesson: robust aggregation mattered because individual models were not equally conservative.

  1. Single-score tri-state decoding: CP082 worked well on the upper tail but classified almost every ordinary image as INDETERMINATE when used for the lower decision.

Lesson: PROTECTED and NOT_PROTECTED should be treated as different statistical problems.

The signal itself remained globally defined

Despite the detector improvements, the deeper Stage 1 architectural limitation was not removed. Stage 2 could search for local evidence much more effectively, but that evidence still came from the original globally defined carrier system. The system therefore worked roughly like:

Global signal
      ↓
crop or reframe
      ↓
surviving fragments of global signal
      ↓
search many possible geometries
      ↓
recover local evidence

This is different from an encoder where each sufficiently large region is intentionally designed to contain self-sufficient protection evidence. Stage 2 improved the recovery of local evidence. It did not yet make the signal itself intrinsically local.


What Stage 2 taught us

  1. The first major lesson was that the Stage 1 crop failure had been partly mischaracterised. The signal was not simply disappearing. Instead useful protection evidence often survived inside cropped regions, but finding it blindly introduced a difficult statistical search problem**. If cropping had physically destroyed the carrier, the obvious next move would have been to increase robustness or signal strength. Instead, Stage 2 showed that a large part of the problem involved representation, search and evidence aggregation.

  2. The second major lesson was that structured and learned methods were complementary. Pure handcrafted correlation had useful inductive structure but insufficient separation. Pure learned replacement discarded information we already knew was meaningful. The strongest Stage 2 detector combined both:

explicit carrier evidence
        +
ordinary-image calibration
        +
learned bounded correction
  1. The third lesson was that false-positive behaviour must be studied in the score tails rather than through average accuracy alone. Several models showed excellent ordering while still producing unacceptable errors around the decision threshold. Training distribution, domain shift and the behaviour of rare difficult ordinary images mattered as much as aggregate model quality.

  2. The fourth lesson was that PROTECTED and NOT_PROTECTED were not naturally symmetric. A detector that is strong at establishing the presence of protection evidence may still be poor at establishing its absence. That led to a more conservative design principle:

No convincing protection evidence
        β‰ 
proof of no protection

The INDETERMINATE state therefore became increasingly important rather than something to minimise at any cost.

  1. Finally, Stage 2 showed that better local decoding alone was unlikely to finish the problem. The detector had become considerably more capable, but the underlying signal still inherited the global geometry of Stage 1. The next architectural question was therefore: Can the signal itself be redesigned so that surviving local regions carry independently useful evidence from the beginning?

Outcome

STAGE 2
COMPLETE_WITH_PROVISIONAL_TRISTATE_THRESHOLD

Established:
βœ“ Cropped regions retain useful local carrier evidence
βœ“ Blind local evidence extraction is possible
βœ“ Shared null normalisation works across geometry hypotheses
βœ“ 32 Γ— 32 overlapping evidence improves the local representation
βœ“ Baseline-preserving residual learning works
βœ“ Tail-weighted training is stable across five seeds
βœ“ Median ensembling improves detector stability
βœ“ Frozen binary detector transfers to Flowers and Wikimedia
βœ“ Separate PROTECTED and NOT_PROTECTED branches are useful
βœ“ Hybrid tri-state architecture produces no observed branch conflicts

Frozen:
βœ“ CP058 single-model lower branch
βœ“ CP078 training recipe
βœ“ CP082 five-model median PROTECTED detector
βœ“ CP085 hybrid tri-state architecture
βœ“ Wikimedia 500-image benchmark manifest

Provisional:
β–³ NOT_PROTECTED threshold policy

Unresolved:
β–³ High INDETERMINATE rate
β–³ Condition-level lower-threshold transfer
β–³ Five-model inference complexity
β–³ Full real-world platform coverage
β–³ Spatial evidence interactions
βœ• Intrinsically local signal representation
βœ• Locally self-sufficient encoding
βœ• Full crop/reframe independence
βœ• Deliberate removal robustness
βœ• Generative transformation robustness

The Stage 2 binary detector was frozen as CP082, the five-model median residual-correction ensemble. The hybrid CP085 structure was also retained as the Stage 2 research architecture:

CP058 lower branch
        +
CP082 upper branch
        +
INDETERMINATE between them

However, the lower threshold remained provisional because one Wikimedia condition reached a 1.60% false-NOT_PROTECTED rate against the intended 1% condition-level limit. No further optimisation should be performed against the already examined Flowers or Wikimedia populations.


Checkpoint record

The detailed development record contains 45 Stage 2 checkpoints. The list below preserves the research trail without requiring a separate page for every experiment.

View all Stage 2 checkpoints
CheckpointFocus
CP044Local evidence baseline and spatial evidence mapping
CP045Blind local alignment hypothesis bank
CP046Null-normalised hypothesis evidence accumulation
CP047Shared null model and one blind threshold
CP048Learned local evidence accumulator
CP049Tiny-set overfit test
CP050Baseline-preserving residual evidence accumulator
CP051Frozen external evaluation and correction audit
CP052Residual correction-cap ablation
CP053Frozen full-VOC comparison
CP054Extended correction-cap boundary
CP055Multi-seed residual-cap stability
CP056Multiscale local-evidence diagnostic
CP057Frozen 32-pixel local baseline transfer
CP058Window-32 residual accumulator
CP059Window-32 residual multi-seed stability
CP060Frozen external window-32 evaluation
CP061Frozen full Oxford Flowers confirmation
CP062Frozen platform-transform stress test
CP063Platform threshold-drift decomposition
CP064Platform-aware shared threshold calibration
CP065External transfer of the platform-aware threshold
CP066Cross-domain false-positive decomposition
CP067Positive-correction support-gate diagnostic
CP068Hypothesis-consensus false-positive audit
CP069Operating-point transition audit
CP070Borderline low-prominence abstention diagnostic
CP071Joint-confidence abstention diagnostic
CP072Learned positive-decision confidence head
CP073Hard-negative confidence training
CP074Asymmetric ordinary correction penalty
CP075Matched-FPR and calibration-drift diagnostic
CP076Score-scale and false-positive mechanism audit
CP077Conditional correction drift audit
CP078High-baseline-tail correction penalty
CP079Tail-weighted multiseed stability
CP080Tail-weighted multiseed external transfer
CP081Five-seed mean-correction ensemble
CP082Robust correction aggregation
CP083Full Flowers fixed-threshold confirmation
CP084Tri-state calibration
CP085Hybrid tri-state decoder
CP086End-to-end evidence-density maps
CP087Hostile internet benchmark
CP088Formal Stage 2 review and artifact freeze

Next steps

Stage 3 should begin from the frozen Stage 2 detector rather than continuing to extract marginal improvements from the same global carrier. The central research question becomes whether GAPP can move from recovering fragments of a globally defined signal to an architecture where local regions are intentionally designed to carry useful protection evidence independently. The next stage should therefore investigate:

  • locally self-sufficient signal representations;
  • overlapping and redundant local encoding;
  • stronger crop and reframing tolerance;
  • whether the existing detector can remain useful as a frozen baseline;
  • perceptual quality of the new representation;
  • real screenshot and platform-processing channels;
  • generative preprocessing and reconstruction;
  • deliberate signal removal and spoofing;
  • the computational cost of blind search;
  • more rigorous tri-state decision semantics.

The Stage 2 architecture should remain available as the reference system while those questions are investigated.