# TopoExploit

A 122M parameter complex-valued autoregressive language model trained on
ExploitGym for security tasks: vulnerability analysis, patch analysis, and
exploit development. It inherits the TopoGPT-2 spectral architecture and is
instrumented with spectral and geometric diagnostics over training dynamics.

This repository contains the model definition, the ExploitGym curriculum
trainer (with EWC + replay anti-forgetting), a small-to-large weight
transfer tool, and an inference engine tailored to exploit-related tasks.

The work is documented in detail in `topogpt3.md`.

## Documentation

- [Quick Start](quickstart.md) — Get running in under five minutes.
- [Tutorial](tutorial.md) — Step-by-step guide from installation to custom training.
- [Essential Concepts](essentials.md) — Core ideas behind complex-valued spectral operators, Grassmannian diagnostics, and HRM.
- [Command Cheatsheet](cheatsheet.md) — Quick reference for CLI commands and Python API.
- [Comparison](comparison.md) — How TopoExploit relates to similar small-scale and security-focused models.
- [Claude Integration Guide](claude.md) — Using TopoExploit with Anthropic models and hybrid pipelines.
- [Technical Paper](topogpt3.md) — Full experimental write-up and results.

## Motivation

Most security and code models scale through size. TopoExploit explores the
opposite direction: whether better representations can let a much smaller
model learn programming and vulnerability structure efficiently. Source code
and binary patches carry strong internal structure (recursion, composition,
scope, repeated motifs), and complex-valued parameters may encode phase
relationships that capture this structure more compactly than real-valued
weights of equal count.

By training on real-world vulnerabilities through the ExploitGym benchmark,
the model aims to reason over C/C++ userspace programs, the V8 engine, and
the Linux kernel — producing vulnerability analyses, patch explanations, and
exploit strategies.

## Checkpoints

Latest trained weights live under `checkpoints_topoexploit/last/` and are
stored as safetensors plus an optimizer file and a JSON state.

## Architecture summary

- Autoregressive transformer with complex-valued spectral operators.
- Quaternion-inspired layers for parameter efficiency.
- A Gauss-style optimization for complex multiplication: three real multiplications per contraction instead of four.
- Sliding window attention (`ATTN_WINDOW`) with configurable window size to control KV-cache memory footprint in long sequences. Set to 0 for full attention or a positive integer for local attention with O(W * S) complexity.
- Latent memory tokens (`N_MEMORY_TOKENS`) for fixed-size context compression across long sequence segments.
- Approximately 122M parameters at the `large` scale.

The base architecture lives in `topogpt3/model.py`. The ExploitGym curriculum
trainer and the Grassmannian / Fisher / phase diagnostics live in
`topogpt3/train.py`.

## Training pipeline

Training proceeds through a three-tier curriculum over the ExploitGym
dataset:

1. Vulnerability analysis (sequence length 256, 6 epochs)
2. Patch analysis (sequence length 512, 3 epochs)
3. Exploit development (sequence length 768, 4 epochs)

Each tier maintains disjoint train, validation and holdout splits. The
holdout is never used during training; it is reserved to measure true
generalization at the end of each tier and at the end of the full pipeline.

### Anti-forgetting: EWC + Replay

Because the tiers are trained sequentially on different task distributions,
the trainer supports Elastic Weight Consolidation (EWC) and replay buffers to
mitigate catastrophic forgetting:

- `--ewc-lambda` (default 5000) — strength of the EWC penalty on important parameters.
- `--replay-ratio` (default 0.25) — fraction of replay samples mixed into each batch.
- `--fisher-batches` (default 50) — number of batches used to estimate the Fisher information.

This is the recommended retraining pass when continuing from an earlier
checkpoint.

### Weight transfer

A small (24.5M) TopoGPT-2 checkpoint can seed a large (122M) TopoExploit
model. `transfer_weights.py` copies matching layers directly, pads
embedding and linear projections, copies matching spectral frequency bins,
and randomly initializes the rest:

```
python transfer_weights.py \
    --from checkpoints_topogpt3/last/model.safetensors \
    --to checkpoints_topoexploit/last/model.safetensors \
    --scale-from small --scale-to large
```

Checkpoints are written atomically to `checkpoints_topoexploit/last/`.
Older `step_*` directories are still loadable for backwards compatibility.

## Optimization diagnostics

At regular intervals the trainer extracts the kernel tensor, performs a
truncated SVD on the leading 16 modes, normalizes them, and records:

- accumulated phase between consecutive normalized dominant kernels,
- net angular drift `W` (a winding-like proxy),
- empirical Fisher spectral gap `Delta_F = lambda_r - lambda_{r+1}`,
- dominant rank `r` from an elbow rule on the singular values.

Grassmannian tracking history is written to `checkpoints_topoexploit/grass_history.jsonl`
and training metrics to `topoexploit_history.jsonl`.

## Inference

The primary inference entry point for exploit tasks is `infer_exploitgym.py`,
which loads the 122M model and generates from prompts derived from the
ExploitGym repository. It supports three tiers:
`vulnerability_analysis`, `patch_analysis`, and `exploit_development`.

Single prompt:

```
python infer_exploitgym.py --prompt "Analyze this vulnerability:" \
    --tier vulnerability_analysis --device cuda
```

Evaluate on holdout tasks:

```
python infer_exploitgym.py --eval-holdout --tier exploit_development \
    --max-samples 50 --device cuda
```

Interactive prompt REPL:

```
python infer_exploitgym.py --interactive --device cuda
```

You can also use the standard autoregressive and hierarchical recursive
reasoning (HRM) samplers from the underlying package.

## One-shot scripts

`run_exploitgym.sh` drives the full prepare / transfer / train / eval
pipeline:

```
./run_exploitgym.sh prepare     -- clone repo + tokenize data
./run_exploitgym.sh transfer    -- transfer weights from small model
./run_exploitgym.sh train       -- train from scratch
./run_exploitgym.sh train-init  -- train with transferred weights
./run_exploitgym.sh eval        -- evaluate on holdout
./run_exploitgym.sh full        -- prepare + transfer + train + eval
```

`run_exploitgym_v2.sh` runs the EWC + replay retraining pass:

```
./run_exploitgym_v2.sh v2-full     -- train v2 with EWC + replay, then eval
./run_exploitgym_v2.sh v2-train    -- train only
./run_exploitgym_v2.sh v2-infer    -- interactive inference
./run_exploitgym_v2.sh v2-eval     -- holdout eval on all three tiers
```

## Repository layout

```
.
├── topogpt3/                  pip-installable package
│   ├── __init__.py            public API re-exports
│   ├── model.py               base TopoGPT2 architecture, tokenizer, helpers
│   ├── train.py               ExploitGym curriculum trainer + diagnostics + EWC/replay
│   ├── exploitgym_config.py   TopoExploit training configuration
│   ├── exploitgym_loader.py   ExploitGym data loader
│   ├── ewc.py                 Elastic Weight Consolidation helpers
│   ├── inference.py           standard autoregressive sampler
│   ├── inference_hrm.py       hierarchical recursive reasoning sampler
│   ├── continuation.py        auto-continuation engine (detect + resume truncated output)
│   ├── lens_model.py          Jacobian-lens model adapter (LensModel protocol)
│   ├── jlens.py               Jacobian lens fitting + application pipeline
│   └── api_server.py          OpenAI-compatible HTTP API server (hardened)
├── tests/                     BDD test suite
├── eval/                      HumanEval benchmark harness
│   ├── harness.py             evaluation pipeline
│   ├── samplers.py            sampler registry (standard / HRM)
│   ├── sandbox.py             sandboxed test executor
│   ├── analysis.py            pass@k / metrics reporting
│   └── diag_static.py         static checkpoint diagnostics
├── infer_exploitgym.py        exploit-task inference entry point
├── transfer_weights.py        small-to-large weight transfer tool
├── run_exploitgym.sh          prepare / transfer / train / eval pipeline
├── run_exploitgym_v2.sh       EWC + replay retraining pipeline
├── app.py                     example entry point for downstream projects
├── Makefile                   convenience targets for all common commands
├── pyproject.toml             package metadata, dependencies, console scripts
├── README.md                  this file
├── topogpt3.md                full paper write-up
├── quickstart.md              five-minute getting started guide
├── tutorial.md                step-by-step usage tutorial
├── essentials.md              core concepts explained
├── cheatsheet.md              command and API quick reference
├── comparison.md              comparison with similar models
├── claude.md                  integration guide for Claude and Anthropic
├── synthetic_dataset.py       optional synthetic dataset helper
├── docs/                      HTML documentation and assets
└── workflows/                 GitHub Actions workflows
```

## Requirements

- Python 3.10 or newer
- PyTorch with CUDA recommended (CPU works for small scales)
- `safetensors`
- `tiktoken` (BPE tokenizer)
- `numpy`
- `datasets` and `huggingface-hub` for data preparation (optional extra `[train]`)

## Installation

From a checkout of this repository:

```
pip install -e .
```

Extra dependencies:

```
pip install -e ".[train]"   # datasets, huggingface-hub
pip install -e ".[lens]"    # huggingface-hub
pip install -e ".[api]"     # fastapi, uvicorn (for the agent harness)
pip install -e ".[dev]"     # pytest, ruff
```

Or install everything at once:

```
pip install -e ".[train,lens,api,dev]"
```

## Command-line usage

Prepare and tokenize ExploitGym data:

```
python -m topogpt3.train --exploitgym --prepare-exploitgym --force-prepare
```

Run the full ExploitGym curriculum:

```
python -m topogpt3.train --exploitgym --train
```

Train with EWC + replay anti-forgetting:

```
python -m topogpt3.train --exploitgym --train --from-scratch \
    --ewc-lambda 5000 --replay-ratio 0.25 --fisher-batches 50
```

Evaluate on the combined holdout:

```
python -m topogpt3.train --exploitgym --eval-holdout
```

Run exploit-task inference:

```
python infer_exploitgym.py --interactive --device cuda
```

### Makefile

A `Makefile` at the repo root wraps common tasks. Run `make help` to see all
targets.

```
make install           pip install -e ".[train,lens,api,dev]"
make test              run full test suite
make lint              ruff check + format
make train             full curriculum training
make infer             standard inference (prompt=def fibonacci)
make infer-continue    inference with auto-continuation
make infer-hrm         HRM inference
make infer-think       HRM thinking mode with auto-continuation
make jlens             Jacobian lens demo (fit 4 prompts)
make api               start API server on port 8800 (no auth)
make api-auth          start API server with authentication
make eval              HumanEval benchmark (all 164 problems)
make eval-sample       HumanEval single problem
make pi                clone, build, and configure Pi agent
make pi-setup          write provider config to ~/.pi/agent/models.json
make pi-run            launch Pi pointed at local API
make clean             remove __pycache__ and .pyc files
```

## Checkpoint compatibility

The model is always built with the maximum sequence length across all
curriculum tiers, so positional embeddings keep a fixed shape regardless of
which tier is used as the entry point. Existing safetensors weights load
without shape mismatch when restarting at a different tier.

As of the 2026-08 sliding window update, the RoPE (Rotary Position Embedding)
caches are stored as non-persistent buffers under the names `_cos_cache` and
`_sin_cache`. Checkpoints from earlier versions that contain `cos_cache` and
`sin_cache` are safely ignored during loading with `strict=False`. The caches
are recomputed to match the current `MAX_SEQ_LEN` configuration at model
instantiation time. Memory tokens (`memory_tokens`) are a new parameter
introduced alongside the latent compression feature; missing this key in
older checkpoints is harmless and produces a warning during load.

## Limitations

This is an exploratory small-scale study. The model is 122M parameters and is
trained on a limited curriculum derived from the ExploitGym benchmark. The
phase and angular drift measurements are diagnostics, not rigorous
mathematical invariants. A real-valued control of the same parameter count,
broader benchmarks, and longer training are needed before drawing stronger
conclusions.

Early generations show syntactic continuity and local semantic consistency.
Exploit correctness and strong algorithmic reasoning remain limited at this
scale and training duration.

## Related work

A 25M-parameter Transformer implementation designed to study language
acquisition as a condensed matter phenomenon. Unlike traditional LLMs,
TopoGPT-2 is engineered to reach a Topological Insulator state — a phase where
grammatical and logical invariants are protected by a spectral gap. Using the
Tiny Stories corpus:

- [https://github.com/grisuno/TopoGPT2](https://github.com/grisuno/TopoGPT2)

The ExploitGym benchmark used for training:

- [https://github.com/ExploitGym](https://github.com/ExploitGym)

## Citation

If you build on this work, please cite:

- [https://doi.org/10.5281/zenodo.20388757](https://doi.org/10.5281/zenodo.20388757)

```
grisuno, "TopoExploit: Exploring Complex-Valued Representations in Small
Security Models", 2026.
```

## License

AGPL v3.
