πŸ› οΈ Current status: Stage 4 | Viability boundaries and protocol hardening Β· Updated 24 August 2026
Stage 3: Local self-sufficiency & protocol definition

The question to address for this stage:

Can GAPP move from recovering fragments of a globally defined signal to a locally self-sufficient system with explicit protocol semantics, defensible tri-state decisions, and clearly measured real-world and adversarial limits?

Stage 2 established that useful GAPP evidence could survive inside cropped image regions. It also showed that recovering that evidence through blind search was expensive and statistically difficult because ordinary image content could accidentally resemble the signal. Stage 3 therefore reopened the underlying representation. Rather than continuing to improve the Stage 2 detector around the same globally defined carrier, we investigated whether the image itself could be encoded with overlapping, locally self-contained regions that remain useful after cropping and reframing. The scope of the research also expanded substantially. Stage 3 did not focus only on the signal and detector. It defined what a GAPP signal actually means, separated PROTECTED, NOT_PROTECTED and INDETERMINATE into explicit decision states, evaluated real codecs and screenshots, tested generative reconstruction, characterised several adversarial weaknesses, and developed the first independently validated negative-clearance rule.

By the end of the stage, GAPP had become a more complete research system. It had also become much clearer where the current system does not work.


Starting point

Stage 3 inherited the frozen Stage 2 system:

Stage 1 global multi-carrier signal
        +
content-adaptive encoder
        +
Stage 2 local evidence search
        +
CP082 five-model PROTECTED detector
        +
CP085 hybrid tri-state architecture

Stage 2 had demonstrated three important facts.

  1. local evidence survived cropping.
  2. a blind detector could recover some of that evidence without the original image, but searching many positions, scales and alignments introduced a substantial ordinary-image false-match tail.
  3. the signal itself was still globally defined.

The detector had become increasingly sophisticated at recovering fragments of a representation that had never been designed to be locally autonomous. Stage 3 therefore began with a different principle:

Do not immediately build a more complicated decoder around the same representation. First make the representation itself local.

The frozen Stage 2 system remained available as a reference rather than being continuously retuned. Stage 3 also began by addressing a separate problem that had become increasingly important: the technical detector could establish evidence of a GAPP signal, but it could not establish who applied it or whether that person had authority to do so. Before changing the architecture, the meaning and threat model therefore had to be defined.


What we explored

1. We first defined what a GAPP signal actually means

The first checkpoints were specification rather than modelling work. The semantic statement frozen at CP089 was:

A GAPP signal asserts that the received image copy carries a machine-readable request that participating generative AI systems do not transform that copy.

This was deliberately defined as a copy-level transformation restriction, not authenticated owner consent. A detected GAPP signal does not establish:

  • ownership;
  • authorship;
  • identity;
  • provenance;
  • authenticity;
  • copyright status;
  • licensing status;
  • legal entitlement;
  • training-data permission;
  • whether the person applying the signal was authorised.

The restriction applies to the received copy. A different unshielded copy of the same underlying work does not automatically inherit it. This distinction became important later when unauthorised shielding and signal transplantation were tested directly.

2. We separated cooperative, accidental and adversarial conditions

Stage 3 also introduced a formal threat model. The research began distinguishing four broad environments:

Cooperative use
        ↓
Participating systems attempt to detect and respect GAPP

Accidental degradation
        ↓
Ordinary processing may weaken the signal

Unauthorised application
        ↓
Someone applies or copies GAPP evidence without authority

Deliberate circumvention
        ↓
Someone intentionally attempts to suppress or bypass detection

These were no longer treated as one generic concept of "robustness". A system surviving JPEG compression does not imply that it resists a motivated attacker. Likewise, a successful adversarial removal does not necessarily mean the mechanism is useless in a cooperative deployment. The claims needed to remain separate.

3. We made the detector states operationally distinct

Stage 3 retained three detector states but clarified their meaning.:

`PROTECTED` means sufficient evidence exists that the received copy carries a GAPP restriction.

`NOT_PROTECTED` requires sufficient evidence for a negative decision.

`INDETERMINATE` means neither one-sided decision is supported strongly enough.

Most importantly, failure to establish PROTECTED does not equate to being NOT_PROTECTED. INDETERMINATE therefore became an explicit escalation state rather than a hidden synonym for allow or refuse. Unsupported detector execution and implementation errors were also kept separate from these evidence states.


4. We tested whether local regions could become independently useful

The main representation branch began by taking the existing signal family and instantiating it independently inside local regions. At 256 Γ— 256 evaluation resolution, a 128 Γ— 128 design created four independent regions. When those local boundaries were known, the result was extremely strong. Across tested conditions:

Local tile pair ordering:
99.8–100%

Images with positive evidence in all four local regions:
~99.4–100%

Positive evidence across all 12 region Γ— RGB components:
~98.8–99.8%

The result also survived JPEG and the tested resize conditions. This established that a local region could carry independently interpretable evidence. But the first design used four non-overlapping quadrants. That created another geometry problem. A centred or off-grid crop could cut through every quadrant and leave zero complete codewords. The representation was locally independent but not sufficiently redundant.

5. Overlapping local redundancy fixed the geometry problem

The next architecture placed 128 Γ— 128 local code regions every 64 pixels. Conceptually:

256 Γ— 256 image

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”     β”‚
β”‚ β”‚ 128   β”‚     β”‚
β”‚ β”‚       β”‚β”Œβ”€β”€β”€β”€β”€
β”‚ β””β”€β”€β”€β”¬β”€β”€β”€β”˜β”‚128 β”‚
β”‚     β”‚    β”‚    β”‚
β”‚ β”Œβ”€β”€β”€β”΄β”€β”€β”€β”β””β”€β”€β”€β”€β”€
β”‚ β”‚ 128   β”‚     β”‚
β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

overlapping local evidence
stride = 64

This fixed the failure exposed by the non-overlapping design. For the tested centre-90%, centre-75% and off-grid-82% crops, the non-overlapping representation left zero complete local codewords. The overlapping design retained at least one complete central codeword. Corner crops retained several. The strongest structured variant became:

overlap_128_s64_sqrt

The square-root overlap rule preserved roughly the earlier visual-quality level while providing substantially stronger redundant local evidence.

6. Blind search was still the harder problem

Solving local geometry did not automatically solve blind detection. Once the detector searched hundreds of possible local position and scale hypotheses, ordinary image content again produced extreme accidental matches. Known alignment could produce strong detection, but blind search substantially raised the decision threshold. One early comparison showed:

Known alignment: 87.4% protected detection

Blind maximum search: 44.8% protected detection

even before difficult cropping was considered. We then tested increasingly structured handcrafted solutions:

  • finer spatial phase search;
  • three-carrier agreement;
  • explicit phase signatures;
  • whole-codeword sign modulation;
  • a dedicated synchronisation carrier;
  • an independent synchronisation pilot.

None solved the core problem. The conclusion we derived to eventually was: The main difficulty was not simply locating the codeword. Exhaustive handcrafted correlation could not reliably distinguish genuine local evidence from the extreme natural-image match tail created by blind search. That closed the handcrafted correlation branch as the primary detector-development path.


7. A small learned local interpreter worked better than more handcrafted rules

The signal remained structured, but the interpretation of each candidate local region became learned. CP106 introduced a very small local CNN with only 47,505 parameters. Patch-level training accuracy increased over six epochs:

68.3%
  ↓
91.5%

The stronger result was conceptual: a learned local interpreter could recognise structured GAPP evidence that simple correlation struggled to distinguish from natural-image matches. However, blind-search calibration remained difficult. A later common calibration gave approximately:

ConditionProtected detectionFalse positives
Clean28.4%0.8%
JPEG23.0%0.6%
Centre crop 90%5.0%0.6%
Centre crop 75%8.2%0.2%
Off-grid crop8.4%0.6%

Pair ordering under the crop conditions nevertheless remained around 95–96%. The representation contained useful information. The detector still needed to learn which of the many blind-search candidates represented genuine local evidence.

8. Aggressive hard-negative mining collapsed the model

The first attempt to solve the ordinary-image search tail was too aggressive. Once online hard-negative mining began, training collapsed toward binary chance:

Loss:      ~0.693
Accuracy:   ~50%
TPR:       0.4–0.6%

This branch was rejected rather than tuned indefinitely. A diagnostic then showed that the model had another problem before hard-negative mining: the actual surviving cropped codeword was often ranked poorly among the blind-search hypotheses. For one warm-up model:

Centre crop 90%
median true-hypothesis rank: 50.5

Centre crop 75%
median rank: 13.5

Off-grid crop
median rank: 32

The next intervention therefore targeted positive training geometry, not merely stronger negative mining.

9. Training on the same hypotheses used at inference fixed localisation

The previous model had been trained on continuous, approximately aligned positive crops. At inference, however, it only received candidates from a discrete blind-search bank. Stage 3 aligned those two geometries. After training positives using the same search hypotheses the detector would actually encounter at runtime, the genuine local codeword moved dramatically upward in the rankings. Typical median ranks became approximately:

Clean        2
JPEG         2
Centre 90    5
Centre 75    1
Off-grid     2

Pair ordering remained around 96–97%. Localisation was largely fixed. The remaining bottleneck was the extreme ordinary-image score tail.

10. Soft hard-negative training preserved localisation

Instead of replacing ordinary negatives with aggressively mined examples, the next curriculum kept the positive-search supervision intact and added only a weak auxiliary hard-negative ranking objective. This worked. At held-out ordinary-image FPRs of only 0.2–0.6%:

ConditionProtected detection
Clean18.8%
JPEG 7514.8%
Centre crop 90% + JPEG21.0%
Centre crop 75% + JPEG17.6%
Off-grid crop + JPEG9.6%

The important result was that localisation remained almost unchanged. The curriculum was then repeated across five random seeds. All five trained without the earlier collapse. The training method, rather than any one selected model, became the Stage 3 learned-local reference.

11. The learned-local result transferred to new images

The five frozen models were first evaluated on 200 Stage-3-unseen images without retraining or threshold refitting. That passed, but those images still came from an older project dataset. A genuinely new domain was therefore opened next. The frozen detector family was evaluated on 500 previously unused Oxford-IIIT Pet test images. No retraining or threshold refitting occurred. All five models transferred successfully, and the new domain performed better than the preceding Caltech evaluation. This provided the first genuine cross-domain evidence for the new local architecture.


12. We tested real codecs rather than only synthetic approximations

Stage 3 then moved into actual byte-level image-processing channels. PNG and JPEG were relatively favourable. WebP was not. Under WebP processing, pair ordering often remained around 97–99% and the correct local region usually remained highly ranked, but the absolute protected score fell substantially. For example:

Frozen-threshold TPR

WebP 80:  10.93%
WebP 60:   4.60%

A diagnostic channel-specific recalibration recovered much of the lost recall:

WebP 80:
10.93% β†’ 40.20%

WebP 60:
 4.60% β†’ 20.20%

JPEG β†’ WebP:
 8.67% β†’ 41.07%

Repeated WebP:
 6.27% β†’ 30.67%

Crop β†’ WebP:
17.80% β†’ 60.60%

However, the result was not purely a threshold shift. At WebP 60, ordinary-image scores shifted by only about βˆ’0.229, while the genuine protected-region score shifted by approximately βˆ’2.227. We therefore characterised WebP as both calibration drift and genuine signal attenuation. A codec control confirmed that the problem was unusually specific to WebP-like processing rather than generic lossy compression. Real JPEG remained substantially stronger even at quality 60.

13. A constrained learned encoder did not improve the system

After identifying the WebP weakness, Stage 3 briefly opened a learned-encoder branch. The experiment remained deliberately constrained. The local carrier itself was not learned. Instead, the encoder learned only how to distribute the fixed embedding budget between nine local regions. The learned weights barely moved from uniform:

~0.981 to 1.016

and the resulting gains were small and inconsistent. The experiment therefore did not earn promotion. Stage 3 retained the structured encoder, overlap_128_s64_sqrt, and the learned CP110 interpretation curriculum.


14. Natural textures exposed genuine signal collisions

A new frozen Describable Textures Dataset (DTD) corpus was then introduced. The result exposed an important failure. On clean DTD images:

Mean ordinary-image FPR:  1.87%
Worst seed:                2.34%
Pair ordering:             ~94%

The increased false-PROTECTED rate was not simply random threshold noise. A follow-up audit found:

11 / 470 images
triggered at least one detector

7 / 470
triggered all five detector seeds

The banded texture class accounted for a disproportionate number of failures, and some images exceeded the protected threshold by very large margins. The conclusion was that some natural visual structures genuinely resemble the current GAPP evidence representation. Stage 3 did not retune the detector around those cases. DTD remained a stress population rather than being silently absorbed into calibration.

15. Ordinary crop and resize geometry was no longer the main weakness

The same DTD corpus was also subjected to geometry-changing benign transformations. The result was almost the reverse of Stage 1. Resize had little effect:

Clean:       ~29.5% TPR
Resize 50%:  ~29.5% TPR

The tested crops also retained useful detection:

Centre crop 90%:  38.04%
Centre crop 75%:  33.96%
Off-grid crop:    30.60%

Rasterisation and mild colour transformations were also relatively well tolerated. The main benign-channel weaknesses had therefore shifted away from crop geometry toward:

natural-pattern collisions
+
WebP attenuation

16. Generative reconstruction caused major signal loss

Stage 3 then tested processing that can occur around generative systems. A real VAE encode/decode reconstruction caused:

Protected detection
29.53% β†’ 7.06%

False positives
 1.87% β†’ 1.96%

Pair ordering
93.96% β†’ 87.02%

Median true-region rank
2 β†’ 3

The ordinary-image tail barely moved while protected evidence weakened substantially. This indicated genuine attenuation through the latent bottleneck rather than a simple calibration shift. Stochastic VAE posterior sampling produced almost the same result. The VAE bottleneck itself, rather than sampling randomness, was therefore the dominant source of loss. A real low-strength image-to-image reconstruction produced another substantial failure. On the full 470-image DTD benchmark at strength 0.15:

Protected detection
29.53% β†’ 6.30%

False positives
 1.87% β†’ 2.17%

Pair ordering
93.96% β†’ 86.47%

Median true-region rank
2 β†’ 4

This established generative reconstruction as a major unsupported channel for the current Stage 3 system.


17. Real screenshots weakened scores but often preserved local structure

Stage 3 also moved beyond synthetic screenshot approximations. A native image-only macOS screenshot benchmark used physical screenshots of matched ordinary/protected image pairs. On the final corrected image-only capture set:

Protected detection
38.33% β†’ 17.50%

Pair ordering
87.50% β†’ 92.50%

Median true-region rank
3.0 β†’ 2.9

The absolute score weakened substantially, but the local evidence structure largely survived. A harder contextual screenshot condition, where the image occupied only part of a larger screenshot, produced much worse localisation. That distinction was retained rather than combining both tests into one result. A physical full-screen iPhone screenshot test also showed that GAPP evidence could still be recovered when the protected image occupied only part of the received screenshot. The remaining problem was again largely the statistical cost of blind search and score attenuation, rather than complete signal destruction.


18. We deliberately tested how the system could be removed or abused

Stage 3 then opened an adversarial suite. The goal was not to prove invulnerability. It was to discover which weaknesses were already practical enough to matter. A frozen corpus of 72 sources was used: 48 Oxford natural images, and 24 DTD texture-stress images. And a fixed family of attacks was evaluated without changing the detector.

Blind stripping

Ordinary image transformations could suppress existing threshold-crossing GAPP evidence. The result was scoped carefully because the natural-image adversarial subset already had low absolute protected detection. The supported conclusion was not that every shield could be reliably removed. It was that some already-detectable protected copies could lose their PROTECTED decision through ordinary quality-preserving transformations.

Signal transplantation

Protected local regions were copied into ordinary target images. The ordinary-patch control produced no new false-PROTECTED decisions. Protected patches produced:

Patch sizeNew false-PROTECTED
64 Γ— 640.00%
96 Γ— 962.08%
128 Γ— 1284.17%
160 Γ— 1608.33%

At the selected 128 Γ— 128 local-region scale, the protected donor patch produced a higher score than the matched ordinary donor patch in 95.83% of cases. The transplant was visually disruptive, so this was not characterised as a stealthy attack But it demonstrated an important consequence of local self-sufficiency:

Useful local GAPP evidence can also be moved.

Unauthorised shielding

Because the Stage 3 encoder contains no authentication mechanism, anyone able to run it can technically apply the signal. The experiment measured whether that actually caused new refusal decisions. Most natural images remained below threshold, but several persisted across multiple independently trained detectors after unauthorised shielding. Two images triggered all five detector seeds. The conclusion was deliberately narrow:

An unauthorised party can create new PROTECTED decisions on some images because the base signal does not authenticate who applied it. This established an authorisation gap rather than universal denial-of-service capability.

Black-box detector probing

We also tested whether someone could adapt to detector feedback. No new false-PROTECTED decisions were synthesised from ordinary images using the tested transform pool. But suppression of existing PROTECTED decisions was highly effective. Edits selected while probing only one detector transferred to held-out detectors:

4 queries:
88.75% suppression
of eligible held-out cases

18 queries:
100% suppression
of eligible held-out cases

Providing the raw scalar detector score did not materially improve threshold-flip success over coarse state feedback. That meant simply hiding the raw score would not remove the underlying weakness.

19. Some adversarial questions remained deliberately unresolved

Unfortunately, stage 3 did not complete every planned adversarial attack. False-evidence synthesis and stronger white-box targeted suppression remained unresolved. They were explicitly recorded as NOT TESTED / DEFERRED, rather than being counted as successful robustness results. This also produced a permanent methodological change.

Longer experiments were required to:

  • run an end-to-end smoke test first;
  • save results incrementally;
  • support resumption;
  • reuse existing score banks where possible;
  • use bounded frozen subsets when full sweeps were unnecessary.

The available research hardware we were using became an explicit experimental constraint rather than something to ignore.


20. We rebuilt the NOT_PROTECTED decision independently

After the robustness and adversarial work, Stage 3 returned to the decision layer. The first evidence inventory found that the new system had substantial positive-evidence data but no independently validated negative-clearance test. Rather than assuming:

not PROTECTED
      ↓
NOT_PROTECTED

the lower tail was studied separately. An offline analysis of existing score banks showed strong separation on natural photographs:

Oxford natural images

AUC:             ~0.987
Pair ordering:   100%

but much poorer separation on texture-heavy DTD content:

DTD

AUC:             ~0.691
Pair ordering:   ~75%

This ruled out treating one universal lower threshold as already justified. The first negative-clearance profile was therefore narrowed to:

DIRECT_NATURAL_PHOTO_V1

Textures, screenshots, generative preprocessing, unknown processing and adversarially suppressed evidence were kept outside that clearance profile.

21. The lower-tail calibration did not transfer by absolute score alone

A fresh replication on VOC natural photographs still showed strong ordering:

AUC:             0.9463
Pair ordering:   100%

but the absolute lower-tail position shifted enough that the earlier finite-sample boundary did not transfer reliably. The shift affected ordinary and protected scores in roughly the same direction. That suggested domain-dependent score origin drift, rather than a collapse in the underlying paired separation. A dedicated negative detector was therefore not added. Instead, Stage 3 retained the existing scalar evidence score and searched for a more defensible semantic boundary.

22. A pre-existing zero boundary became the N candidate

The normalized evidence coordinate had a useful pre-existing reference:

score = 0

Zero corresponded to the original detector's logit boundary. It had therefore not been invented after inspecting the final validation set. The candidate rule was pre-registered:

Three-seed median normalized score

score < 0
        ↓
NOT_PROTECTED

using detector seeds:

4110
4112
4114

A fresh untouched validation set was then frozen before the rule was evaluated. It contained:

96 Oxford natural-photo sources
96 VOC natural-photo sources
192 total

No post-validation threshold adjustment was permitted.


What we built

The final Stage 3 research architecture was substantially different from Stage 2.

Signal and encoder

Structured local representation:
overlap_128_s64_sqrt

Local codeword:
128 Γ— 128

Stride:
64

Property:
overlapping regions carry locally useful evidence

The Stage 1 global carrier remained only as an earlier reference. The constrained learned strength allocator was not promoted.

Local detector

The selected learned-local training reference used:

Search-aligned positive training
        +
low-weight auxiliary hard-negative ranking
        +
blind local hypothesis search

The training curriculum reproduced across five random seeds. The five models were used as evidence of reproducibility rather than automatically becoming a deployment ensemble.

Negative-clearance profile

The Stage 3 negative decision was explicitly scoped:

Profile:
DIRECT_NATURAL_PHOTO_V1

Statistic:
median normalized score
across seeds 4110 / 4112 / 4114

Rule:
score < 0.0

This rule did not apply automatically to:

  • textures;
  • screenshots;
  • generative reconstruction;
  • unknown preprocessing;
  • adversarially suppressed evidence.

Final tri-state integration

For the supported direct natural-photo profile:

score < 0.0
        ↓
NOT_PROTECTED

0.0 ≀ score < 1.0
        ↓
INDETERMINATE

score β‰₯ 1.0
        ↓
PROTECTED

1.0 was the normalized form of the already frozen positive detector thresholds. It was not a new threshold fitted against the final validation set.


Key results

Untouched NOT_PROTECTED validation

The pre-registered negative rule was evaluated exactly once on the untouched 192-source natural-photo validation set.

DatasetProtected false NOT_PROTECTEDOrdinary clearedOne-sided 95% upper bound
Oxford0 / 9651.04%3.07%
VOC0 / 9650.00%3.07%
Pooled0 / 19250.52%1.55%

No protected validation copy was incorrectly cleared as NOT_PROTECTED. At the same time, roughly half of ordinary natural-photo copies received a definitive negative decision. This was sufficient to promote the rule within the tested profile, not universally.

Final integrated tri-state validation

The final direct natural-photo audit produced:

192 protected copies

PROTECTED:        43
INDETERMINATE:   149
NOT_PROTECTED:     0

and:

192 ordinary copies

NOT_PROTECTED:    97
INDETERMINATE:    94
PROTECTED:         1

The integration therefore produced:

Protected false NOT_PROTECTED:  0
P / N conflicts:                 0
Out-of-profile silent N:         0

The result was intentionally conservative. Only:

43 / 192
= 22.4%

of the protected validation copies received a definitive PROTECTED decision. That is an important limitation of the Stage 3 system. It should not be directly compared with the Stage 1 or Stage 2 detection percentages as if these were identical experiments: the representation, datasets, evaluation roles and final decision procedure were different. The takeaway is simply that positive definitive coverage remained weak and required further investigation.

Benign-channel summary

The local representation handled several ordinary geometric operations comparatively well:

Resize
        β†’ generally stable

Tested crop/reframe geometry
        β†’ local evidence survives

JPEG / PNG
        β†’ comparatively favourable

Raster / mild colour changes
        β†’ comparatively favourable

The two strongest benign-content limitations were:

WebP attenuation
        +
natural texture collisions

Generative-channel summary

Generative reconstruction remained a major weakness:

VAE reconstruction
29.53% β†’ 7.06% TPR

Low-strength img2img
29.53% β†’ 6.30% TPR

The signal was weakened substantially before detection.

Adversarial summary

Stage 3 established that the current system is not adversarially secure. Evidence supported:

βœ“ Blind stripping can suppress some existing P decisions

βœ“ Protected evidence can be transplanted

βœ“ Unauthorised shielding can create new P decisions on some images

βœ“ Black-box probing can find transferable suppression transforms

while:

? False-evidence synthesis
? White-box targeted suppression

remained unresolved rather than assumed safe.


What worked

Several Stage 3 ideas were strong enough to retain.

  1. Local codewords proved technically viable: Independent 128 Γ— 128 regions carried extremely strong aligned evidence, establishing that local self-sufficiency was possible in principle.

  2. Overlapping redundancy solved the core crop geometry problem: A stride-64 layout allowed intact local codewords to survive centre and off-grid crops that destroyed every codeword in the non-overlapping design.

  3. Learned local interpretation outperformed continued handcrafted search rules: A small neural interpreter could distinguish structured local evidence more effectively than increasingly complicated correlation and synchronisation heuristics.

  4. Matching training geometry to inference geometry fixed localisation: Training on the same discrete hypothesis bank used at runtime moved genuine cropped codewords from poor ranks to approximately 5 / 1 / 2 across the tested crop conditions.

  5. Soft hard-negative ranking controlled the ordinary tail without destroying the representation: The CP110 curriculum improved low-FPR detection while preserving the localisation breakthrough.

  6. The training recipe reproduced across five seeds: Stage 3 demonstrated that the learned-local behaviour was not dependent on one fortunate initialization.

  7. The local architecture transferred to a genuinely new image domain: The frozen detector family transferred to 500 previously unused Oxford-IIIT Pet images without retraining or threshold refitting.

  8. Ordinary crop and resize robustness improved substantially: On the new local representation, arbitrary crop geometry was no longer the dominant failure seen in Stage 1.

  9. Real screenshots preserved meaningful local structure: Native macOS and iPhone captures weakened absolute scores, but useful local evidence could still survive the screenshot pipeline.

  10. PROTECTED and NOT_PROTECTED remained independent decisions: The final decision layer did not interpret weak positive evidence as permission.

  11. A pre-registered negative-clearance rule passed untouched validation: The scoped natural-photo rule produced 0 / 192 protected false clearances while definitively clearing approximately half of ordinary copies.

  12. Adversarial weaknesses were measured rather than hidden: Removal, transplantation, unauthorised application and black-box probing were retained as explicit limitations of the Stage 3 system.


What did not work

Stage 3 also closed or rejected several major branches.

  1. Non-overlapping local regions: A four-quadrant 128 Γ— 128 design left zero intact codewords after several centred and off-grid crops.

Lesson: local independence without spatial redundancy is not sufficient.

  1. Finer blind phase search: Reducing spatial mismatch substantially did not materially improve detection.

Lesson: location resolution was not the main source of the search-tail problem.

  1. Stronger three-carrier consistency rules: These reduced some ordinary-image tail behaviour but did not recover the lost blind-search separation.

Lesson: genuine evidence cannot be isolated by simple carrier agreement alone.

  1. Whole-codeword synchronisation masks: Explicit sign modulation reduced payload detectability and slightly worsened quality.

Lesson: synchronisation cannot be added by destructively modulating the existing payload.

  1. Dedicated synchronisation carrier: Reserving one carrier for sync did not solve blind crop detection.

Lesson: the main failure was deeper than lack of an explicit phase marker.

  1. Independent handcrafted synchronisation pilot: The pilot did not overcome natural-image match tails.

Lesson: further template-correlation engineering was unlikely to be the productive path.

  1. Aggressive online hard-negative mining: Training collapsed to near chance.

Lesson: hard negatives must remain auxiliary rather than replacing the positive/ordinary training distribution.

  1. Constrained learned strength allocation: Reweighting the existing nine local regions barely changed the encoder and did not consistently improve WebP robustness.

Lesson: simple redistribution of a fixed signal budget was not enough.

  1. Universal texture handling: DTD exposed persistent natural-pattern collisions across multiple detector seeds.

Lesson: current detector calibration cannot automatically be assumed to transfer to every visual content class.

  1. WebP robustness: Local structure often survived, but frozen-threshold evidence was materially attenuated.

Lesson: some codecs change both calibration and the actual signal strength.

  1. VAE reconstruction: Protected detection fell from 29.53% to 7.06%.

Lesson: the latent reconstruction bottleneck can substantially erase usable evidence.

  1. Low-strength image-to-image reconstruction: Protected detection fell to 6.30% on the full benchmark.

Lesson: even relatively mild generative reconstruction is outside the demonstrated robustness envelope.

  1. Universal NOT_PROTECTED calibration: A lower-tail rule discovered on one natural domain did not transfer by absolute score position to another.

Lesson: negative clearance requires profile-scoped calibration and independent validation.

  1. Adversarial stripping resistance: Some transformations could suppress existing threshold-crossing evidence.

Lesson: the current signal cannot be treated as removal-resistant.

  1. Transplantation resistance: Large protected regions could move detector evidence into another image.

Lesson: local self-sufficiency creates a corresponding copying surface.

  1. Authorisation: The base signal could be applied by parties whose authority the detector cannot verify.

Lesson: signal presence and legitimate authority are different technical problems.

  1. Resistance to black-box probing: Coarse detector feedback was sufficient to discover highly transferable suppression transforms.

Lesson: hiding raw scores does not remove the current detector-family weakness.

  1. Full adversarial coverage: False-evidence synthesis and white-box targeted suppression were not completed.

Lesson: untested attack classes must remain explicitly unresolved rather than being interpreted as robustness.

Positive definitive coverage remained weak

The final tri-state integration was safe on the tested negative-clearance population, but conservative. Only 43 / 192 protected copies received a definitive PROTECTED result. Another 149 / 192 remained INDETERMINATE. This means Stage 3 did not establish that the current architecture has a sufficiently large positive decision region for practical deployment. That question remained open at closeout.


What Stage 3 taught us

Stage 3 answered the original representation question positively, but with significant qualifications. The global-coordinate failure from Stage 1 was not fundamental.

  1. A structured overlapping local representation can place independently useful evidence throughout an image, and that evidence can survive cropping without reconstructing the original complete frame. The main geometry problem therefore shifted. The question was no longer:

Can a surviving crop contain a complete signal? It became: Can a blind detector distinguish that surviving local signal from the strongest accidental matches produced by ordinary image content?

The learned-local branch showed that this is possible to a meaningful extent, but blind search still creates a difficult statistical tail and considerable computational cost. Stage 3 also established that robustness is not one property. The architecture behaves very differently under these conditions:

JPEG
crop
resize
screenshots
WebP
VAE reconstruction
img2img reconstruction
natural textures
deliberate removal
adaptive probing

These conditions cannot be collapsed into one statement that the signal either "survives" or "does not survive". The supported operating envelope has to be explicit.

  1. A second major lesson was that detection and authority are separate. The detector can estimate whether GAPP evidence exists. It cannot determine whether whoever applied that evidence had the right to do so. Transplantation and unauthorised shielding made this limitation concrete rather than theoretical.

  2. The third major lesson concerned the tri-state system. PROTECTED and NOT_PROTECTED are different one-sided statistical decisions. Strong evidence for one is not simply the inverse of the other. The conservative INDETERMINATE region is therefore a necessary part of the current system rather than an implementation inconvenience.

  3. Finally, Stage 3 showed the value of allowing the research to produce uncomfortable results. The project finished this stage with a more credible local architecture and a validated scoped negative-clearance rule, but also with:

weak positive definitive coverage
natural texture collisions
WebP attenuation
generative reconstruction failure
blind-search cost
removal vulnerability
transplantation
an authorisation gap
and unresolved stronger attacks

Those limitations became part of the result.


Outcome

STAGE 3
COMPLETE_WITH_VALIDATED_PROFILE_SCOPED_TRISTATE_DECISION_LAYER
AND_KNOWN_ADVERSARIAL_LIMITATIONS

Established:
βœ“ Copy-level GAPP semantics defined
βœ“ Threat model defined
βœ“ PROTECTED / NOT_PROTECTED / INDETERMINATE behaviour defined
βœ“ Global carrier retired from active development
βœ“ Locally self-sufficient 128 Γ— 128 code regions demonstrated
βœ“ Overlapping stride-64 redundancy solves tested crop geometry
βœ“ Structured overlap_128_s64_sqrt encoder selected
βœ“ Learned local interpretation is viable
βœ“ Search-aligned positive training works
βœ“ Soft hard-negative curriculum works
βœ“ Training recipe reproduced across five seeds
βœ“ Genuine cross-domain learned-local transfer
βœ“ Crop / resize / JPEG / PNG behaviour characterised
βœ“ Native macOS and iPhone screenshot behaviour characterised
βœ“ DIRECT_NATURAL_PHOTO_V1 negative-clearance profile defined
βœ“ Pre-registered N rule validated on untouched data
βœ“ 0 / 192 protected false NOT_PROTECTED decisions
βœ“ One-sided P / N tri-state architecture integrated
βœ“ Blind removal limitations characterised
βœ“ Transplantation vulnerability characterised
βœ“ Unauthorised-shielding gap characterised
βœ“ Black-box suppression vulnerability characterised

Frozen:
βœ“ Structured encoder: overlap_128_s64_sqrt
βœ“ Learned-local training reference: CP110 curriculum
βœ“ N aggregation seeds: 4110 / 4112 / 4114
βœ“ DIRECT_NATURAL_PHOTO_V1
βœ“ NOT_PROTECTED rule: normalized median score < 0.0
βœ“ Tri-state integration architecture

Profile-scoped:
β–³ NOT_PROTECTED validation applies to direct natural photographs
β–³ Textures do not inherit the N rule
β–³ Screenshots do not inherit the N rule
β–³ Generative preprocessing does not inherit the N rule
β–³ Unknown preprocessing does not inherit the N rule
β–³ Adversarially suppressed evidence does not inherit the N rule

Unresolved:
β–³ Positive definitive coverage: 43 / 192
β–³ Large protected INDETERMINATE region: 149 / 192
β–³ Blind-search computational cost
β–³ Natural texture collisions
β–³ WebP attenuation
β–³ Screenshot score attenuation
βœ• VAE reconstruction robustness
βœ• Low-strength img2img robustness
βœ• Robustness to deliberate stripping
βœ• Transplantation resistance
βœ• Authenticated authority
βœ• Resistance to black-box suppression
? False-evidence synthesis
? White-box targeted suppression

The formal Stage 3 review approved the stage for closeout. The principal research artifacts carried forward were:

Structured local encoder
overlap_128_s64_sqrt

Learned-local training reference
CP110 curriculum

Scoped negative-clearance profile
DIRECT_NATURAL_PHOTO_V1

Negative rule
median normalized score < 0.0

Integrated direct-photo decision bands
< 0.0       β†’ NOT_PROTECTED
0.0 to <1.0 β†’ INDETERMINATE
β‰₯ 1.0       β†’ PROTECTED

The architecture should not be interpreted as universal, production-ready or adversarially secure. Stage 3 closes with a validated profile-scoped decision layer and a substantially clearer map of the system's operating boundaries.


Checkpoint record

The detailed development record contains 84 Stage 3 checkpoints. The list below preserves the research trail without requiring a separate page for every experiment.

View all Stage 3 checkpoints
CheckpointFocus
CP089Define GAPP permission semantics
CP090Define the GAPP threat model
CP091Define detector-state behaviour and escalation
CP092Define the operational error model
CP093Define the perceptual-quality protocol
CP094Audit the frozen Stage 2 baseline
CP095Test independent local codewords
CP096Test aligned local-codeword feasibility
CP097Test blind local phase-and-scale search
CP098Decompose the blind-search failure
CP099Introduce overlapping local redundancy
CP100Test overlapping codewords under blind detection
CP101Sweep blind spatial phase resolution
CP102Test three-carrier consistency
CP103Add an explicit local synchronisation signature
CP104Separate synchronisation and payload carriers
CP105Test an independent synchronisation signal
CP106Introduce a learned local interpreter
CP107Train with crop and online hard negatives
CP108Diagnose hard-negative training collapse
CP109Align positive training with inference search
CP110Add soft hard-negative fine-tuning
CP111Replicate the learned-local curriculum across seeds
CP112Test frozen internal holdout transfer
CP113Test a genuinely new benign image domain
CP114Test real WebP processing
CP115Separate WebP calibration drift from attenuation
CP116Compare WebP with real JPEG and PNG
CP117Test a constrained learned strength allocator
CP118Select the Stage 3 encoder and training architecture
CP119Define the real-channel benchmark plan
CP120Freeze the DTD benign-channel corpus
CP121Run the DTD real-codec benchmark
CP122Audit natural-texture false-PROTECTED cases
CP123Test geometry-changing benign channels
CP124Test rasterisation and colour channels
CP125Close reproducible benign Suite A
CP126Test deterministic VAE reconstruction
CP127Test stochastic VAE posterior sampling
CP128Pilot low-strength real img2img reconstruction
CP129Run the full low-strength img2img benchmark
CP130Close generative-processing Suite B
CP131Freeze the native macOS screenshot protocol
CP132Generate the macOS screenshot capture bundle
CP133Perform the native macOS captures
CP134Validate native macOS captures
CP135Evaluate native macOS screenshots
CP136Freeze the native mobile screenshot protocol
CP137Validate the native mobile capture set
CP138Evaluate native mobile screenshots
CP139Close the real-platform screenshot study
CP140Define the adversarial evaluation suite
CP141Freeze the adversarial corpus and attack set
CP142Run blind removal and stripping tests
CP143Close the blind-removal threat family
CP144Define the copying/transplantation protocol
CP145Run the copying/transplantation benchmark
CP146Close the transplantation threat family
CP147Define the unauthorised-shielding test
CP148Run the unauthorised-shielding benchmark
CP149Close the unauthorised-shielding threat family
CP150Define the black-box probing protocol
CP151Run the black-box probing benchmark
CP152Close the black-box probing threat family
CP153Define the false-evidence synthesis protocol
CP154Attempt the bounded false-evidence evaluation
CP155Close adversarial Suite C and freeze compute policy
CP156Define the Stage 3 tri-state decision protocol
CP157Inventory positive and negative decision evidence
CP158Audit offline lower-tail separability
CP159Freeze the negative-clearance calibration protocol
CP160Freeze a fresh lower-tail replication corpus
CP161Run negative-clearance replication
CP162Decompose lower-tail transfer failure
CP163Freeze the negative-clearance architecture decision
CP164Freeze the multi-domain calibration corpus
CP165Score the multi-domain calibration corpus
CP166Audit the semantic zero-score boundary
CP167Pre-register the candidate NOT_PROTECTED rule
CP168Freeze the untouched validation corpus
CP169Run one-shot untouched N validation
CP170Promote the scoped NOT_PROTECTED rule
CP171Audit final tri-state integration
CP172Formal Stage 3 review and closeout

Next steps

Stage 4 should begin from the frozen Stage 3 architecture rather than assuming that another encoder redesign is automatically required. The most important unresolved question is now practical viability:

Does the current image-native, database-free GAPP design possess a useful decision region under clearly defined cooperative, lossy and adversarial conditions?

The first priority should be to understand why only 43 / 192 protected validation copies reached the final PROTECTED state. Possible causes include:

positive threshold
        ↓
blind search
        ↓
learned interpretation
        ↓
signal representation
        ↓
or a combination of them

This should be diagnosed before changing the architecture. Stage 4 should also:

  • map the full positive-recall / false-PROTECTED operating frontier;
  • stratify weak protected cases by image properties and evidence structure;
  • determine how much recall is lost specifically to blind search;
  • evaluate genuinely new external image populations;
  • compare GAPP directly with credible blind image-marker baselines;
  • separate cooperative channels from generative laundering and active adversarial attacks;
  • revisit WebP, VAE and generative reconstruction as distinct channel classes;
  • quantify blind-search runtime and implementation cost;
  • keep PROTECTED and NOT_PROTECTED as independent one-sided decisions;
  • preserve the authorisation boundary between a base GAPP signal and any future authenticated layer;
  • accept a narrower cooperative protocol or a rigorous negative result if the present constraints do not support a useful operating region.

The Stage 3 system should remain frozen as the reference while those questions are answered.