# Polyglot Codebase Knowledge Graph

> Generated offline by **readmenator**. 44 files, 824 symbols, 414 imports. Supports C, C++, Python, Go, Rust, JS/TS, Java, C#, Shell, PHP, Dart, GDScript, Nim, ASM, Ruby, Swift, Kotlin, Scala, Lua, Elixir.
> No LLMs. No tokens. Pure static analysis. See more [here](https://github.com/grisuno/ReadMenator)

**Start here:** Statistics Dashboard for scope, God Nodes for blast radius, Architecture Reference for per-file API. Agents: prefer `readmenator-agent/INDEX.md` + `SYMBOLS.md`.

**Wiki:** prefer `readmenator-wiki/index.md` for progressive disclosure: one synthesis page per community, `connections.json` with EXTRACTED vs INFERRED confidence, `queries.md` log, `REPORT.md` audit.

**Confidence:** EXTRACTED = parsed from source, INFERRED = heuristic bridge, AMBIGUOUS = reported, never hidden. See `readmenator-wiki/REPORT.md`.

**Total Files Parsed:** 44 | **Total Symbols Extracted:** 824 | **Total Imports:** 414
 | **Resolved Imports:** 66

<!-- ranking_model: v1.0 | weights: {ppr:0.45,auth:0.2,test:0.15,doc:0.1,fresh:0.1} | alpha:0.85 | commit:82e48da | date:2026-07-18 -->


## Table of Contents

1. [Statistics Dashboard](#statistics-dashboard)
2. [Architectural Layers](#architectural-layers)
3. [Ranked Context](#ranked-context)
4. [God Nodes](#god-nodes)
5. [Community Analysis](#community-analysis)
6. [Surprising Connections](#surprising-connections)
7. [Suggested Questions](#suggested-questions)
8. [Force-Graph Explorer](#force-graph-explorer)
9. [Corpus Analytics](#corpus-analytics)
10. [Taint Propagation Map](#taint-propagation-map)
11. [Hotspot Analysis](#hotspot-analysis)
12. [Change Impact Analysis](#change-impact-analysis)
13. [Suggested Linting Rules](#suggested-linting-rules)
14. [Dataflow Analysis](#dataflow-analysis)
15. [Concept Graph](#concept-graph)
16. [Orphans](#orphans)
17. [Query Recipes](#query-recipes)
18. [Structural Knowledge Map](#structural-knowledge-map)
19. [UML Class Diagram](#uml-class-diagram)
20. [Code Property Graph](#code-property-graph)
21. [Architecture Reference](#architecture-reference)
    - [PY (39 files)](#py-39-files)
    - [SH (5 files)](#sh-5-files)

---

## Statistics Dashboard

| Metric | Value |
|--------|-------|
| Total Files | 44 |
| Total Symbols | 824 |
| Total Imports | 414 |
| Call Edges | 5593 |
| Inheritance Edges | 36 |
| Languages | 2 |
| Avg Symbols/File | 18.7 |
| Avg Imports/File | 9.4 |
| Resolved Imports | 66 |

### Top Files by Import Count (Fan-Out)

| File | Imports | Symbols | Language |
|------|---------|---------|----------|
| `api_server.py` | 27 | 46 | py |
| `train.py` | 26 | 64 | py |
| `model.py` | 25 | 187 | py |
| `harness.py` | 24 | 12 | py |
| `synthetic_dataset.py` | 24 | 35 | py |
| `repair.py` | 16 | 6 | py |
| `noise_sweep.py` | 14 | 4 | py |
| `jlens.py` | 14 | 29 | py |
| `lens_model.py` | 14 | 29 | py |
| `diag_static.py` | 13 | 5 | py |

---

## Architectural Layers

Auto-detected from path patterns, naming conventions, and imported frameworks.

| Layer | Files |
|-------|-------|
| utility | 36 |
| testing | 2 |
| infrastructure | 2 |
| business_logic | 2 |
| data_access | 1 |
| presentation | 1 |

### utility

- `app.py` (py, 5 symbols)
- `analyze.py` (py, 5 symbols)
- `analyze_results.py` (py, 4 symbols)
- `diag_static.py` (py, 5 symbols)
- `governor.py` (py, 20 symbols)
- `governor_smoke.py` (py, 7 symbols)
- `harness.py` (py, 12 symbols)
- `hodge_cm_ablation.py` (py, 6 symbols)
- `integration_smoke.py` (py, 1 symbols)
- `noise_analysis.py` (py, 3 symbols)
- `noise_sweep.py` (py, 4 symbols)
- `repair.py` (py, 6 symbols)
- `report.py` (py, 6 symbols)
- `samplers.py` (py, 7 symbols)
- `sandbox.py` (py, 9 symbols)
- *... and 21 more*

### data_access

- `synthetic_dataset.py` (py, 35 symbols)

### testing

- `test_jlens.py` (py, 51 symbols)
- `test_lens_model.py` (py, 34 symbols)

### presentation

- `api_server.py` (py, 46 symbols)

### infrastructure

- `exploitgym_config.py` (py, 3 symbols)
- `merged_config.py` (py, 6 symbols)

### business_logic

- `lens_model.py` (py, 29 symbols)
- `model.py` (py, 187 symbols)

---

## Ranked Context

Files ranked by composite score for the current query context. The ranking combines Personalized PageRank (query relevance), global authority, test coverage, documentation coverage, and code freshness. Model: v1.0.

| Rank | File | Composite | PPR | Authority | Test | Doc |
|------|------|-----------|-----|-----------|------|-----|
| 1 | `model.py` | 0.1531 | 0.1607 | 0.1607 | 0.00 | 0.49 |
| 2 | `continuation.py` | 0.1483 | 0.1051 | 0.1051 | 0.00 | 0.80 |
| 3 | `app.py` | 0.1275 | 0.0116 | 0.0116 | 0.00 | 1.20 |
| 4 | `ewc.py` | 0.1201 | 0.0169 | 0.0169 | 0.00 | 1.09 |
| 5 | `integration_smoke.py` | 0.1075 | 0.0116 | 0.0116 | 0.00 | 1.00 |
| 6 | `sandbox_smoke.py` | 0.1075 | 0.0116 | 0.0116 | 0.00 | 1.00 |
| 7 | `__main__.py` | 0.1075 | 0.0116 | 0.0116 | 0.00 | 1.00 |
| 8 | `transfer_weights.py` | 0.1075 | 0.0116 | 0.0116 | 0.00 | 1.00 |
| 9 | `__init__.py` | 0.1071 | 0.0815 | 0.0815 | 0.00 | 1.00 |
| 10 | `merged_config.py` | 0.1007 | 0.0268 | 0.0268 | 0.00 | 0.83 |

---

## God Nodes

Most architecturally central files ranked by combined import/export degree and symbol richness.

| File | Score | Connections | PageRank |
|------|-------|-------------|----------|
| `model.py` | 48.7 | | 0.1607 |
| `__init__.py` | 30.0 | | 0.0815 |
| `train.py` | 24.4 | | 0.0000 |
| `lens_model.py` | 14.9 | | 0.0000 |
| `inference_hrm.py` | 13.6 | | 0.0000 |
| `harness.py` | 13.2 | | 0.0000 |
| `jlens.py` | 12.9 | | 0.0000 |
| `__main__.py` | 12.1 | | 0.0116 |
| `api_server.py` | 10.6 | | 0.0000 |
| `exploitgym_loader.py` | 9.9 | | 0.0000 |

---

## Community Analysis

Files grouped by import-based community detection. Cohesion measures how tightly connected each community is internally.

### topogpt3: model (Cohesion: 0.64)

**12 files** in this community:

- `diag_static.py` (py, 5 symbols)
- `hodge_cm_ablation.py` (py, 6 symbols)
- `infer_exploitgym.py` (py, 11 symbols)
- `synthetic_dataset.py` (py, 35 symbols)
- `ewc.py` (py, 11 symbols)
- `exploitgym_config.py` (py, 3 symbols)
- `exploitgym_loader.py` (py, 19 symbols)
- `hodge_cm.py` (py, 15 symbols)
- `merged_config.py` (py, 6 symbols)
- `model.py` (py, 187 symbols)
- `train.py` (py, 64 symbols)
- `transfer_weights.py` (py, 3 symbols)

### eval: harness (Cohesion: 0.64)

**7 files** in this community:

- `harness.py` (py, 12 symbols)
- `integration_smoke.py` (py, 1 symbols)
- `noise_sweep.py` (py, 4 symbols)
- `samplers.py` (py, 7 symbols)
- `sandbox.py` (py, 9 symbols)
- `sandbox_smoke.py` (py, 1 symbols)
- `temp_sweep.py` (py, 5 symbols)

### eval: governor (Cohesion: 0.31)

**6 files** in this community:

- `app.py` (py, 5 symbols)
- `governor.py` (py, 20 symbols)
- `governor_smoke.py` (py, 7 symbols)
- `repair.py` (py, 6 symbols)
- `smoke.py` (py, 2 symbols)
- `__init__.py` (py, 0 symbols)

### topogpt3: inference_hrm (Cohesion: 0.42)

**5 files** in this community:

- `__main__.py` (py, 1 symbols)
- `api_server.py` (py, 46 symbols)
- `continuation.py` (py, 5 symbols)
- `inference.py` (py, 54 symbols)
- `inference_hrm.py` (py, 76 symbols)

### topogpt3: jlens (Cohesion: 0.45)

**4 files** in this community:

- `test_jlens.py` (py, 51 symbols)
- `test_lens_model.py` (py, 34 symbols)
- `jlens.py` (py, 29 symbols)
- `lens_model.py` (py, 29 symbols)

---

## Surprising Connections

Files in different communities connected through 3+ indirect hops.

- `sandbox_smoke.py` <-> `hodge_cm.py` (6 hops, across 2 communities)
- `governor.py` <-> `sandbox_smoke.py` (5 hops, across 2 communities)
- `governor.py` <-> `hodge_cm.py` (5 hops, across 2 communities)
- `hodge_cm_ablation.py` <-> `sandbox_smoke.py` (5 hops, across 2 communities)
- `integration_smoke.py` <-> `hodge_cm.py` (5 hops, across 2 communities)

---

## Suggested Questions

Auto-generated exploration prompts based on graph structure:

- What does model.py depend on, and what depends on it? (15 connections)
- What does __init__.py depend on, and what depends on it? (15 connections)
- What does train.py depend on, and what depends on it? (9 connections)
- How are the 12 files in 'topogpt3: model' related to each other?
- Why are sandbox_smoke.py and hodge_cm.py connected through 6 hops across 2 communities?

---

## Force-Graph Explorer

Heterogeneous explorer payload (files, communities, layers, externals) rendered with force-graph physics, stable family colors, log2 node sizing, and convex-hull community overlays. Open `readmenator-maps/graph-force.html` (linked from the maps gallery) or run `explorer`.

- Nodes: 107 (community: 5, external: 52, file: 44, layer: 6)
- Edges: 492

---

## Corpus Analytics

Attribution funnel and distributions for the explorer dashboard.

- Files: 44 | With symbols: 40 | Attributed: 34 | God nodes: 10 | Symbols: 824

- Layer `utility`: 36 files
- Layer `business_logic`: 2 files
- Layer `infrastructure`: 2 files
- Layer `testing`: 2 files
- Layer `data_access`: 1 files
- Layer `presentation`: 1 files

- Language `py`: 39 files
- Language `sh`: 5 files

---

## Taint Propagation Map

Taint analysis traces how dangerous imports propagate through the codebase via transitive dependencies. Source files import dangerous modules directly; sink files receive the danger indirectly.

**Taint Sources:** 2 | **Taint Sinks:** 18 | **Propagation Paths:** 20

- `governor_smoke.py` imports `eval` (0 hop to `governor_smoke.py`) [critical]
  Path: governor_smoke.py
- `governor_smoke.py` imports `eval` (1 hop to `governor.py`) [critical]
  Path: governor_smoke.py -> governor.py
- `governor_smoke.py` imports `eval` (1 hop to `__init__.py`) [critical]
  Path: governor_smoke.py -> __init__.py
- `governor_smoke.py` imports `eval` (2 hops to `merged_config.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> merged_config.py
- `governor_smoke.py` imports `eval` (2 hops to `inference.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> inference.py
- `governor_smoke.py` imports `eval` (2 hops to `train.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> train.py
- `governor_smoke.py` imports `eval` (2 hops to `jlens.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> jlens.py
- `governor_smoke.py` imports `eval` (2 hops to `lens_model.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> lens_model.py
- `governor_smoke.py` imports `eval` (2 hops to `inference_hrm.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> inference_hrm.py
- `governor_smoke.py` imports `eval` (2 hops to `model.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> model.py
- `governor_smoke.py` imports `eval` (3 hops to `exploitgym_loader.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> merged_config.py -> exploitgym_loader.py
- `governor_smoke.py` imports `eval` (3 hops to `ewc.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> train.py -> ewc.py
- `governor_smoke.py` imports `eval` (3 hops to `exploitgym_config.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> train.py -> exploitgym_config.py
- `governor_smoke.py` imports `eval` (3 hops to `continuation.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> inference_hrm.py -> continuation.py
- `governor_smoke.py` imports `eval` (3 hops to `synthetic_dataset.py`) [critical]
  Path: governor_smoke.py -> __init__.py -> model.py -> synthetic_dataset.py
- `harness.py` imports `subprocess` (0 hop to `harness.py`) [high]
  Path: harness.py
- `harness.py` imports `subprocess` (1 hop to `samplers.py`) [high]
  Path: harness.py -> samplers.py
- `harness.py` imports `subprocess` (1 hop to `sandbox.py`) [high]
  Path: harness.py -> sandbox.py
- `harness.py` imports `subprocess` (1 hop to `__init__.py`) [high]
  Path: harness.py -> __init__.py
- `harness.py` imports `subprocess` (2 hops to `merged_config.py`) [high]
  Path: harness.py -> __init__.py -> merged_config.py

---

## Hotspot Analysis

Files ranked by combined complexity (symbol count) and centrality (connection count). High-scoring files are architecturally critical and may need refactoring attention.

| File | Complexity | Centrality | Combined | Symbols | Connections |
|------|-----------|------------|----------|---------|-------------|
| `model.py` | 1.000 | 1.000 | 1.000 | 187 | 42 |
| `continuation.py` | 0.027 | 0.143 | 0.096 | 5 | 6 |
| `app.py` | 0.027 | 0.167 | 0.111 | 5 | 7 |
| `ewc.py` | 0.059 | 0.167 | 0.123 | 11 | 7 |
| `integration_smoke.py` | 0.005 | 0.095 | 0.059 | 1 | 4 |
| `sandbox_smoke.py` | 0.005 | 0.095 | 0.059 | 1 | 4 |
| `__main__.py` | 0.005 | 0.333 | 0.202 | 1 | 14 |
| `transfer_weights.py` | 0.016 | 0.214 | 0.135 | 3 | 9 |
| `__init__.py` | 0.000 | 0.548 | 0.329 | 0 | 23 |
| `merged_config.py` | 0.032 | 0.214 | 0.141 | 6 | 9 |
| `train.py` | 0.342 | 0.833 | 0.637 | 64 | 35 |
| `api_server.py` | 0.246 | 0.714 | 0.527 | 46 | 30 |
| `harness.py` | 0.064 | 0.738 | 0.469 | 12 | 31 |
| `synthetic_dataset.py` | 0.187 | 0.595 | 0.432 | 35 | 25 |
| `inference_hrm.py` | 0.406 | 0.381 | 0.391 | 76 | 16 |

---

## Dataflow Analysis

Procedural intra-function dataflow findings (zero tokens, regex-based heuristics, all INFERRED). Each lead is grounded at file:line for manual review.

**1 findings** (UNCHECKED_ALLOC: 1).

| File | Function | Line | Kind | Variable | Description |
|------|----------|------|------|----------|-------------|
| `eval/repair.py` | `main` | 140 | `UNCHECKED_ALLOC` | `base` | Result of allocator stored in `base` is never checked against NULL. |

---

## Concept Graph

Semantic second-brain layer: nouns are concept nodes, verbs are edges. Each noun maps atomically to a file set (EXTRACTED); each verb aggregates structural imports, calls, and inherits into consumes, invokes, extends, depends_on, or bridges (INFERRED).

**50 concepts, 100 relations.**

| Concept | Files | Mentions |
|---------|-------|----------|
| `topo` | 27 | 99 |
| `model` | 25 | 186 |
| `eval` | 25 | 77 |
| `run` | 22 | 70 |
| `gpt3` | 20 | 48 |
| `topogpt3` | 20 | 38 |
| `load` | 19 | 40 |
| `build` | 18 | 44 |
| `prompt` | 16 | 64 |
| `returns` | 16 | 53 |
| `when` | 16 | 29 |
| `checkpoint` | 15 | 67 |
| `config` | 15 | 58 |
| `one` | 15 | 27 |
| `all` | 14 | 37 |
| `new` | 14 | 28 |
| `runs` | 14 | 24 |
| `layer` | 13 | 83 |
| `per` | 13 | 52 |
| `pass` | 13 | 37 |
| `code` | 13 | 30 |
| `token` | 12 | 88 |
| `file` | 12 | 57 |
| `return` | 12 | 49 |
| `error` | 12 | 34 |
| `same` | 12 | 15 |
| `max` | 11 | 37 |
| `top` | 11 | 36 |
| `each` | 11 | 30 |
| `train` | 11 | 28 |

### Verb Edges

| Source | Verb | Target | Strength | Evidence |
|--------|------|--------|----------|----------|
| `topo` | `depends_on` | `topogpt3` | 1.00 | 10 |
| `topo` | `depends_on` | `model` | 0.96 | 10 |
| `gpt3` | `depends_on` | `topogpt3` | 0.85 | 10 |
| `gpt3` | `depends_on` | `model` | 0.81 | 10 |
| `gpt3` | `depends_on` | `topo` | 0.81 | 10 |
| `topo` | `depends_on` | `config` | 0.79 | 10 |
| `topo` | `depends_on` | `new` | 0.79 | 10 |
| `model` | `depends_on` | `topogpt3` | 0.77 | 10 |
| `topo` | `depends_on` | `checkpoint` | 0.75 | 10 |
| `topo` | `depends_on` | `returns` | 0.75 | 10 |
| `topo` | `depends_on` | `file` | 0.73 | 10 |
| `model` | `depends_on` | `topo` | 0.71 | 10 |
| `topo` | `depends_on` | `all` | 0.71 | 10 |
| `topo` | `depends_on` | `build` | 0.71 | 10 |
| `topo` | `depends_on` | `prompt` | 0.71 | 10 |
| `topo` | `depends_on` | `run` | 0.71 | 10 |
| `checkpoint` | `depends_on` | `topogpt3` | 0.69 | 10 |
| `gpt3` | `depends_on` | `new` | 0.69 | 10 |
| `topo` | `depends_on` | `when` | 0.69 | 10 |
| `checkpoint` | `depends_on` | `model` | 0.67 | 10 |
| `gpt3` | `depends_on` | `config` | 0.67 | 10 |
| `topo` | `depends_on` | `gpt2` | 0.67 | 10 |
| `topo` | `depends_on` | `layer` | 0.67 | 10 |
| `topo` | `depends_on` | `load` | 0.67 | 10 |
| `checkpoint` | `depends_on` | `topo` | 0.65 | 10 |
| `gpt3` | `depends_on` | `checkpoint` | 0.65 | 10 |
| `topo` | `depends_on` | `max` | 0.65 | 10 |
| `topo` | `depends_on` | `tensor` | 0.65 | 10 |
| `topogpt3` | `depends_on` | `model` | 0.65 | 10 |
| `gpt3` | `depends_on` | `returns` | 0.62 | 10 |

### Dialectic Prompts

- Thesis: `all` centralizes 14 files; Antithesis: `checkpoint` pulls 15 files with 9 shared (Jaccard 0.45); Synthesis: should they merge, split by layer, or keep `bridges` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `config` pulls 15 files with 9 shared (Jaccard 0.45); Synthesis: should they merge, split by layer, or keep `bridges` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `each` pulls 11 files with 6 shared (Jaccard 0.32); Synthesis: should they merge, split by layer, or keep `bridges` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `error` pulls 12 files with 7 shared (Jaccard 0.37); Synthesis: should they merge, split by layer, or keep `bridges` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `file` pulls 12 files with 10 shared (Jaccard 0.62); Synthesis: should they merge, split by layer, or keep `bridges` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `full` pulls 10 files with 6 shared (Jaccard 0.33); Synthesis: should they merge, split by layer, or keep `bridges` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `gpt2` pulls 10 files with 6 shared (Jaccard 0.33); Synthesis: should they merge, split by layer, or keep `bridges` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `gpt3` pulls 20 files with 9 shared (Jaccard 0.36); Synthesis: should they merge, split by layer, or keep `depends_on` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `layer` pulls 13 files with 9 shared (Jaccard 0.50); Synthesis: should they merge, split by layer, or keep `bridges` explicit?
- Thesis: `all` centralizes 14 files; Antithesis: `load` pulls 19 files with 8 shared (Jaccard 0.32); Synthesis: should they merge, split by layer, or keep `bridges` explicit?

---

## Change Impact Analysis

Files sorted by how many other files would be affected if they changed. High-impact files should be changed with caution.

| File | Direct Dependents | Transitive Dependents | Total Impact |
|------|------------------|----------------------|--------------|
| `continuation.py` | 3 | 23 | 26 |
| `synthetic_dataset.py` | 1 | 24 | 25 |
| `model.py` | 13 | 11 | 24 |
| `exploitgym_loader.py` | 3 | 13 | 16 |
| `lens_model.py` | 5 | 10 | 15 |
| `ewc.py` | 1 | 13 | 14 |
| `exploitgym_config.py` | 1 | 13 | 14 |
| `jlens.py` | 4 | 10 | 14 |
| `merged_config.py` | 2 | 12 | 14 |
| `train.py` | 4 | 9 | 13 |
| `inference.py` | 2 | 10 | 12 |
| `inference_hrm.py` | 2 | 10 | 12 |
| `__init__.py` | 8 | 2 | 10 |
| `sandbox.py` | 2 | 3 | 5 |
| `samplers.py` | 1 | 3 | 4 |

---

## Suggested Linting Rules

Automatically suggested linting and security rules based on patterns detected in the codebase. These can be exported as Semgrep rules using the `--export-rules` flag.

| Rule ID | Severity | Description | Language | Matches |
|---------|----------|-------------|----------|---------|
| `RM001` | info | Large number of functions in py: 672 total | py | 672 |
| `RM002` | info | Large number of functions in sh: 11 total | sh | 11 |
| `RM003` | info | Print statement found (consider logging instead) | python | 210 |

---

## Orphans

Files with no documentation or low connectivity. These are candidates for documentation investment or cleanup.

- `install.sh` (0 symbols, no doc)
- `run_exploitgym_v2.sh` (4 symbols, no doc)
- `run_merged.sh` (7 symbols, no doc)

---

## Query Recipes

Example queries you can run against this knowledge base using the ranking engine:

```
# Find files most relevant to a concept
readmenator query "Where is the import resolver implemented?"

# Rank files by relevance to a topic
readmenator query "How does documentation generation work?"

# Explain why a file ranks highly
readmenator query "explain readmenator/_documentation.py"

# Trace dependency paths with ranked context
readmenator query "path from CLI to exporter"
```

The ranking model uses the following signals:

- **Personalized PageRank** (45% weight): query-specific relevance via seed propagation
- **Global Authority** (20% weight): structural importance via standard PageRank
- **Test Coverage** (15% weight): fraction of symbols referenced in test files
- **Doc Coverage** (10% weight): presence of docstrings and file-level docs
- **Freshness** (10% weight): recent modification activity

Results include score decomposition and justification paths for each ranked item.

---

## Structural Knowledge Map

```mermaid
graph TD
    classDef mod fill:#1e1e1e,stroke:#ff6666,stroke-width:2px,color:#fff;
    classDef cls fill:#2d2d2d,stroke:#4ec9b0,stroke-width:2px,color:#fff;
    classDef fn fill:#333,stroke:#dcdcaa,stroke-width:1px,color:#dcdcaa;
    classDef ext fill:#111,stroke:#666,stroke-dasharray:5 5,color:#aaa;
    subgraph community_0 ["topogpt3: model"]
    topogpt3_train_py["train.py (py)"]
    class topogpt3_train_py mod;
    topogpt3_train_py_TopoGPT3Config["TopoGPT3Config"]
    class topogpt3_train_py_TopoGPT3Config cls;
    topogpt3_train_py --> topogpt3_train_py_TopoGPT3Config
    topogpt3_train_py_GrassmannianTracker["GrassmannianTracker"]
    class topogpt3_train_py_GrassmannianTracker cls;
    topogpt3_train_py --> topogpt3_train_py_GrassmannianTracker
    topogpt3_train_py__gauss_complex_contract["_gauss_complex_contract"]
    class topogpt3_train_py__gauss_complex_contract fn;
    topogpt3_train_py --> topogpt3_train_py__gauss_complex_contract
    topogpt3_train_py_apply_gauss_patch["apply_gauss_patch"]
    class topogpt3_train_py_apply_gauss_patch fn;
    topogpt3_train_py --> topogpt3_train_py_apply_gauss_patch
    topogpt3_train_py_EfficiencyMetrics["EfficiencyMetrics"]
    class topogpt3_train_py_EfficiencyMetrics cls;
    topogpt3_train_py --> topogpt3_train_py_EfficiencyMetrics
    end
    subgraph community_3 ["topogpt3: inference_hrm"]
    topogpt3_api_server_py["api_server.py (py)"]
    class topogpt3_api_server_py mod;
    end
    subgraph community_1 ["eval: harness"]
    eval_harness_py["harness.py (py)"]
    class eval_harness_py mod;
    topogpt3_model_py["model.py (py)"]
    class topogpt3_model_py mod;
    synthetic_dataset_py["synthetic_dataset.py (py)"]
    class synthetic_dataset_py mod;
    end
    subgraph community_4 ["topogpt3: jlens"]
    tests_test_lens_model_py["test_lens_model.py (py)"]
    class tests_test_lens_model_py mod;
    end
    subgraph community_2 ["eval: governor"]
    eval_repair_py["repair.py (py)"]
    class eval_repair_py mod;
    eval_noise_sweep_py["noise_sweep.py (py)"]
    class eval_noise_sweep_py mod;
    topogpt3_jlens_py["jlens.py (py)"]
    class topogpt3_jlens_py mod;
    topogpt3_lens_model_py["lens_model.py (py)"]
    class topogpt3_lens_model_py mod;
    eval_diag_static_py["diag_static.py (py)"]
    class eval_diag_static_py mod;
    topogpt3___init___py["__init__.py (py)"]
    class topogpt3___init___py mod;
    topogpt3_inference_hrm_py["inference_hrm.py (py)"]
    class topogpt3_inference_hrm_py mod;
    topogpt3___main___py["__main__.py (py)"]
    class topogpt3___main___py mod;
    infer_exploitgym_py["infer_exploitgym.py (py)"]
    class infer_exploitgym_py mod;
    eval_temp_sweep_py["temp_sweep.py (py)"]
    class eval_temp_sweep_py mod;
    topogpt3_exploitgym_loader_py["exploitgym_loader.py (py)"]
    class topogpt3_exploitgym_loader_py mod;
    eval_hodge_cm_ablation_py["hodge_cm_ablation.py (py)"]
    class eval_hodge_cm_ablation_py mod;
    topogpt3_inference_py["inference.py (py)"]
    class topogpt3_inference_py mod;
    eval_sandbox_py["sandbox.py (py)"]
    class eval_sandbox_py mod;
    eval_governor_smoke_py["governor_smoke.py (py)"]
    class eval_governor_smoke_py mod;
    topogpt3_vanilla_control_py["vanilla_control.py (py)"]
    class topogpt3_vanilla_control_py mod;
    eval_report_py["report.py (py)"]
    class eval_report_py mod;
    eval_noise_analysis_py["noise_analysis.py (py)"]
    class eval_noise_analysis_py mod;
    transfer_weights_py["transfer_weights.py (py)"]
    class transfer_weights_py mod;
    eval_governor_py["governor.py (py)"]
    class eval_governor_py mod;
    eval_analyze_py["analyze.py (py)"]
    class eval_analyze_py mod;
    tests_test_jlens_py["test_jlens.py (py)"]
    class tests_test_jlens_py mod;
    topogpt3_merged_config_py["merged_config.py (py)"]
    class topogpt3_merged_config_py mod;
    app_py["app.py (py)"]
    class app_py mod;
    topogpt3_exploitgym_config_py["exploitgym_config.py (py)"]
    class topogpt3_exploitgym_config_py mod;
    topogpt3_ewc_py["ewc.py (py)"]
    class topogpt3_ewc_py mod;
    eval_samplers_py["samplers.py (py)"]
    class eval_samplers_py mod;
    eval_analyze_results_py["analyze_results.py (py)"]
    class eval_analyze_results_py mod;
    topogpt3_hodge_cm_py["hodge_cm.py (py)"]
    class topogpt3_hodge_cm_py mod;
    eval_smoke_py["smoke.py (py)"]
    class eval_smoke_py mod;
    eval_integration_smoke_py["integration_smoke.py (py)"]
    class eval_integration_smoke_py mod;
    eval_sandbox_smoke_py["sandbox_smoke.py (py)"]
    class eval_sandbox_smoke_py mod;
    topogpt3_continuation_py["continuation.py (py)"]
    class topogpt3_continuation_py mod;
    run_merged_sh["run_merged.sh (sh)"]
    class run_merged_sh mod;
    run_exploitgym_v2_sh["run_exploitgym_v2.sh (sh)"]
    class run_exploitgym_v2_sh mod;
    install_sh["install.sh (sh)"]
    class install_sh mod;
    run_exploitgym_sh["run_exploitgym.sh (sh)"]
    class run_exploitgym_sh mod;
    run_hodge_cm_ablation_sh["run_hodge_cm_ablation.sh (sh)"]
    class run_hodge_cm_ablation_sh mod;
    end
    app_py -- resolved_imports --> topogpt3___init___py
    eval_diag_static_py -- resolved_imports --> topogpt3___init___py
    eval_diag_static_py -- resolved_imports --> topogpt3_model_py
    eval_diag_static_py -- resolved_imports --> topogpt3_train_py
    eval_governor_smoke_py -- resolved_imports --> eval_governor_py
    eval_governor_smoke_py -- resolved_imports --> topogpt3___init___py
    eval_harness_py -- resolved_imports --> topogpt3___init___py
    eval_harness_py -- resolved_imports --> eval_samplers_py
    eval_harness_py -- resolved_imports --> eval_sandbox_py
    eval_harness_py -- resolved_imports --> eval_samplers_py
    eval_hodge_cm_ablation_py -- resolved_imports --> topogpt3_hodge_cm_py
    eval_hodge_cm_ablation_py -- resolved_imports --> topogpt3_model_py
    eval_integration_smoke_py -- resolved_imports --> eval_harness_py
    eval_noise_sweep_py -- resolved_imports --> topogpt3___init___py
    eval_noise_sweep_py -- resolved_imports --> topogpt3_model_py
    eval_noise_sweep_py -- resolved_imports --> eval_harness_py
    eval_repair_py -- resolved_imports --> topogpt3___init___py
    eval_samplers_py -- resolved_imports --> topogpt3___init___py
    eval_sandbox_smoke_py -- resolved_imports --> eval_sandbox_py
    eval_smoke_py -- resolved_imports --> topogpt3___init___py
    eval_temp_sweep_py -- resolved_imports --> eval_noise_sweep_py
    eval_temp_sweep_py -- resolved_imports --> eval_harness_py
    infer_exploitgym_py -- resolved_imports --> topogpt3_model_py
    infer_exploitgym_py -- resolved_imports --> topogpt3_train_py
    tests_test_jlens_py -- resolved_imports --> topogpt3_lens_model_py
    tests_test_jlens_py -- resolved_imports --> topogpt3_jlens_py
    tests_test_lens_model_py -- resolved_imports --> topogpt3_lens_model_py
    tests_test_lens_model_py -- resolved_imports --> topogpt3_model_py
    tests_test_lens_model_py -- resolved_imports --> topogpt3_model_py
    tests_test_lens_model_py -- resolved_imports --> topogpt3_jlens_py
    tests_test_lens_model_py -- resolved_imports --> topogpt3_jlens_py
    tests_test_lens_model_py -- resolved_imports --> topogpt3_jlens_py
    tests_test_lens_model_py -- resolved_imports --> topogpt3_jlens_py
    topogpt3___init___py -- resolved_imports --> topogpt3_model_py
    topogpt3___init___py -- resolved_imports --> topogpt3_train_py
    topogpt3___init___py -- resolved_imports --> topogpt3_merged_config_py
    topogpt3___init___py -- resolved_imports --> topogpt3_inference_py
    topogpt3___init___py -- resolved_imports --> topogpt3_inference_hrm_py
    topogpt3___init___py -- resolved_imports --> topogpt3_lens_model_py
    topogpt3___init___py -- resolved_imports --> topogpt3_jlens_py
    topogpt3___main___py -- resolved_imports --> topogpt3_jlens_py
    topogpt3___main___py -- resolved_imports --> topogpt3_lens_model_py
    topogpt3___main___py -- resolved_imports --> topogpt3_api_server_py
    topogpt3___main___py -- resolved_imports --> topogpt3_inference_py
    topogpt3___main___py -- resolved_imports --> topogpt3_inference_hrm_py
    topogpt3___main___py -- resolved_imports --> topogpt3_train_py
    topogpt3_api_server_py -- resolved_imports --> topogpt3_model_py
    topogpt3_api_server_py -- resolved_imports --> topogpt3_continuation_py
    topogpt3_exploitgym_config_py -- resolved_imports --> topogpt3_model_py
    topogpt3_exploitgym_config_py -- resolved_imports --> topogpt3_exploitgym_loader_py
    topogpt3_exploitgym_loader_py -- resolved_imports --> topogpt3_model_py
    topogpt3_inference_hrm_py -- resolved_imports --> topogpt3_continuation_py
    topogpt3_jlens_py -- resolved_imports --> topogpt3_lens_model_py
    topogpt3_jlens_py -- resolved_imports --> topogpt3_lens_model_py
    topogpt3_lens_model_py -- resolved_imports --> topogpt3_model_py
    topogpt3_lens_model_py -- resolved_imports --> topogpt3_model_py
    topogpt3_merged_config_py -- resolved_imports --> topogpt3_model_py
    topogpt3_merged_config_py -- resolved_imports --> topogpt3_exploitgym_loader_py
    topogpt3_model_py -- resolved_imports --> synthetic_dataset_py
    topogpt3_model_py -- resolved_imports --> topogpt3_continuation_py
    topogpt3_train_py -- resolved_imports --> topogpt3_model_py
    topogpt3_train_py -- resolved_imports --> topogpt3_exploitgym_loader_py
    topogpt3_train_py -- resolved_imports --> topogpt3_exploitgym_config_py
    topogpt3_train_py -- resolved_imports --> topogpt3_merged_config_py
    topogpt3_train_py -- resolved_imports --> topogpt3_ewc_py
    transfer_weights_py -- resolved_imports --> topogpt3_model_py
    ext___future__["__future__"]
    class ext___future__ ext;
    app_py -.->|imports| ext___future__
    ext_argparse["argparse"]
    class ext_argparse ext;
    app_py -.->|imports| ext_argparse
    ext_sys["sys"]
    class ext_sys ext;
    app_py -.->|imports| ext_sys
    ext_typing["typing"]
    class ext_typing ext;
    app_py -.->|imports| ext_typing
    ext_torch["torch"]
    class ext_torch ext;
    app_py -.->|imports| ext_torch
    ext_topogpt3["topogpt3"]
    class ext_topogpt3 ext;
    app_py -.->|imports| ext_topogpt3
    eval_analyze_py -.->|imports| ext___future__
    eval_analyze_py -.->|imports| ext_argparse
    ext_json["json"]
    class ext_json ext;
    eval_analyze_py -.->|imports| ext_json
    ext_math["math"]
    class ext_math ext;
    eval_analyze_py -.->|imports| ext_math
    ext_re["re"]
    class ext_re ext;
    eval_analyze_py -.->|imports| ext_re
    ext_collections["collections"]
    class ext_collections ext;
    eval_analyze_py -.->|imports| ext_collections
    ext_pathlib["pathlib"]
    class ext_pathlib ext;
    eval_analyze_py -.->|imports| ext_pathlib
    eval_analyze_py -.->|imports| ext_typing
    eval_analyze_results_py -.->|imports| ext___future__
    eval_analyze_results_py -.->|imports| ext_argparse
    eval_analyze_results_py -.->|imports| ext_json
    eval_analyze_results_py -.->|imports| ext_collections
    eval_analyze_results_py -.->|imports| ext_pathlib
    eval_diag_static_py -.->|imports| ext___future__
    eval_diag_static_py -.->|imports| ext_argparse
    eval_diag_static_py -.->|imports| ext_json
    eval_diag_static_py -.->|imports| ext_math
    eval_diag_static_py -.->|imports| ext_sys
    ext_time["time"]
    class ext_time ext;
    eval_diag_static_py -.->|imports| ext_time
    eval_diag_static_py -.->|imports| ext_pathlib
    eval_diag_static_py -.->|imports| ext_typing
    eval_diag_static_py -.->|imports| ext_torch
    eval_diag_static_py -.->|imports| ext_topogpt3
    ext_topogpt3_model["topogpt3.model"]
    class ext_topogpt3_model ext;
    eval_diag_static_py -.->|imports| ext_topogpt3_model
    ext_safetensors_torch["safetensors.torch"]
    class ext_safetensors_torch ext;
    eval_diag_static_py -.->|imports| ext_safetensors_torch
    ext_topogpt3_train["topogpt3.train"]
    class ext_topogpt3_train ext;
    eval_diag_static_py -.->|imports| ext_topogpt3_train
    eval_governor_py -.->|imports| ext___future__
    ext_threading["threading"]
    class ext_threading ext;
    eval_governor_py -.->|imports| ext_threading
    eval_governor_py -.->|imports| ext_time
    ext_dataclasses["dataclasses"]
    class ext_dataclasses ext;
    eval_governor_py -.->|imports| ext_dataclasses
    ext_enum["enum"]
    class ext_enum ext;
    eval_governor_py -.->|imports| ext_enum
    eval_governor_py -.->|imports| ext_typing
    eval_governor_py -.->|imports| ext_torch
    ext_torch_nn_functional["torch.nn.functional"]
    class ext_torch_nn_functional ext;
    eval_governor_py -.->|imports| ext_torch_nn_functional
    eval_governor_smoke_py -.->|imports| ext_sys
    eval_governor_smoke_py -.->|imports| ext_threading
    eval_governor_smoke_py -.->|imports| ext_time
    eval_governor_smoke_py -.->|imports| ext_pathlib
    eval_governor_smoke_py -.->|imports| ext_torch
    ext_eval_governor["eval.governor"]
    class ext_eval_governor ext;
    eval_governor_smoke_py -.->|imports| ext_eval_governor
    eval_governor_smoke_py -.->|imports| ext_topogpt3
    ext_safetensors["safetensors"]
    class ext_safetensors ext;
    eval_governor_smoke_py -.->|imports| ext_safetensors
    eval_governor_smoke_py -.->|imports| ext_safetensors_torch
    eval_harness_py -.->|imports| ext___future__
    eval_harness_py -.->|imports| ext_argparse
    ext_contextlib["contextlib"]
    class ext_contextlib ext;
    eval_harness_py -.->|imports| ext_contextlib
    ext_io["io"]
    class ext_io ext;
    eval_harness_py -.->|imports| ext_io
    eval_harness_py -.->|imports| ext_json
    ext_os["os"]
    class ext_os ext;
    eval_harness_py -.->|imports| ext_os
    eval_harness_py -.->|imports| ext_re
    ext_signal["signal"]
    class ext_signal ext;
    eval_harness_py -.->|imports| ext_signal
    ext_subprocess["subprocess"]
    class ext_subprocess ext;
    eval_harness_py -.->|imports| ext_subprocess
    eval_harness_py -.->|imports| ext_sys
    eval_harness_py -.->|imports| ext_time
    ext_traceback["traceback"]
    class ext_traceback ext;
    eval_harness_py -.->|imports| ext_traceback
    eval_harness_py -.->|imports| ext_dataclasses
    eval_harness_py -.->|imports| ext_pathlib
    eval_harness_py -.->|imports| ext_typing
    eval_harness_py -.->|imports| ext_torch
    eval_harness_py -.->|imports| ext_torch
    eval_harness_py -.->|imports| ext_topogpt3
    ext_eval_samplers["eval.samplers"]
    class ext_eval_samplers ext;
    eval_harness_py -.->|imports| ext_eval_samplers
    eval_harness_py -.->|imports| ext_safetensors_torch
    ext_datasets["datasets"]
    class ext_datasets ext;
    eval_harness_py -.->|imports| ext_datasets
    ext_eval_sandbox["eval.sandbox"]
    class ext_eval_sandbox ext;
    eval_harness_py -.->|imports| ext_eval_sandbox
    eval_harness_py -.->|imports| ext_eval_samplers
    eval_harness_py -.->|imports| ext_safetensors
    eval_hodge_cm_ablation_py -.->|imports| ext_argparse
    eval_hodge_cm_ablation_py -.->|imports| ext_json
    eval_hodge_cm_ablation_py -.->|imports| ext_math
    eval_hodge_cm_ablation_py -.->|imports| ext_sys
    eval_hodge_cm_ablation_py -.->|imports| ext_pathlib
    eval_hodge_cm_ablation_py -.->|imports| ext_torch
    ext_numpy["numpy"]
    class ext_numpy ext;
    eval_hodge_cm_ablation_py -.->|imports| ext_numpy
    eval_hodge_cm_ablation_py -.->|imports| ext_safetensors_torch
    ext_topogpt3_hodge_cm["topogpt3.hodge_cm"]
    class ext_topogpt3_hodge_cm ext;
    eval_hodge_cm_ablation_py -.->|imports| ext_topogpt3_hodge_cm
    eval_hodge_cm_ablation_py -.->|imports| ext_topogpt3_model
    eval_integration_smoke_py -.->|imports| ext_sys
    eval_integration_smoke_py -.->|imports| ext_pathlib
    ext_eval_harness["eval.harness"]
    class ext_eval_harness ext;
    eval_integration_smoke_py -.->|imports| ext_eval_harness
    eval_noise_analysis_py -.->|imports| ext___future__
    eval_noise_analysis_py -.->|imports| ext_argparse
    ext_ast["ast"]
    class ext_ast ext;
    eval_noise_analysis_py -.->|imports| ext_ast
    eval_noise_analysis_py -.->|imports| ext_json
    eval_noise_analysis_py -.->|imports| ext_re
    eval_noise_analysis_py -.->|imports| ext_sys
    eval_noise_analysis_py -.->|imports| ext_collections
    eval_noise_analysis_py -.->|imports| ext_pathlib
    eval_noise_analysis_py -.->|imports| ext_typing
    eval_noise_sweep_py -.->|imports| ext___future__
    eval_noise_sweep_py -.->|imports| ext_argparse
    eval_noise_sweep_py -.->|imports| ext_json
    eval_noise_sweep_py -.->|imports| ext_math
    eval_noise_sweep_py -.->|imports| ext_sys
    eval_noise_sweep_py -.->|imports| ext_time
    eval_noise_sweep_py -.->|imports| ext_pathlib
    eval_noise_sweep_py -.->|imports| ext_typing
    eval_noise_sweep_py -.->|imports| ext_torch
    eval_noise_sweep_py -.->|imports| ext_topogpt3
    eval_noise_sweep_py -.->|imports| ext_topogpt3_model
    eval_noise_sweep_py -.->|imports| ext_safetensors_torch
    eval_noise_sweep_py -.->|imports| ext_safetensors
    eval_noise_sweep_py -.->|imports| ext_eval_harness
    eval_repair_py -.->|imports| ext___future__
    eval_repair_py -.->|imports| ext_argparse
    eval_repair_py -.->|imports| ext_contextlib
    eval_repair_py -.->|imports| ext_io
    eval_repair_py -.->|imports| ext_json
    eval_repair_py -.->|imports| ext_re
    eval_repair_py -.->|imports| ext_time
    eval_repair_py -.->|imports| ext_traceback
    eval_repair_py -.->|imports| ext_collections
    eval_repair_py -.->|imports| ext_pathlib
    eval_repair_py -.->|imports| ext_typing
    eval_repair_py -.->|imports| ext_torch
    eval_repair_py -.->|imports| ext_safetensors
    eval_repair_py -.->|imports| ext_safetensors_torch
    eval_repair_py -.->|imports| ext_topogpt3
    eval_repair_py -.->|imports| ext_datasets
    eval_report_py -.->|imports| ext___future__
    eval_report_py -.->|imports| ext_argparse
    eval_report_py -.->|imports| ext_json
    eval_report_py -.->|imports| ext_math
    eval_report_py -.->|imports| ext_re
    ext_shutil["shutil"]
    class ext_shutil ext;
    eval_report_py -.->|imports| ext_shutil
    ext_statistics["statistics"]
    class ext_statistics ext;
    eval_report_py -.->|imports| ext_statistics
    eval_report_py -.->|imports| ext_collections
    eval_report_py -.->|imports| ext_pathlib
    eval_samplers_py -.->|imports| ext___future__
    eval_samplers_py -.->|imports| ext_os
    eval_samplers_py -.->|imports| ext_typing
    eval_samplers_py -.->|imports| ext_topogpt3
    eval_sandbox_py -.->|imports| ext___future__
    eval_sandbox_py -.->|imports| ext_ast
    eval_sandbox_py -.->|imports| ext_os
    eval_sandbox_py -.->|imports| ext_subprocess
    eval_sandbox_py -.->|imports| ext_sys
    ext_tempfile["tempfile"]
    class ext_tempfile ext;
    eval_sandbox_py -.->|imports| ext_tempfile
    ext_textwrap["textwrap"]
    class ext_textwrap ext;
    eval_sandbox_py -.->|imports| ext_textwrap
    eval_sandbox_py -.->|imports| ext_dataclasses
    eval_sandbox_py -.->|imports| ext_pathlib
    eval_sandbox_py -.->|imports| ext_typing
    eval_sandbox_py -.->|imports| ext_json
    eval_sandbox_smoke_py -.->|imports| ext_sys
    eval_sandbox_smoke_py -.->|imports| ext_pathlib
    eval_sandbox_smoke_py -.->|imports| ext_eval_sandbox
    eval_smoke_py -.->|imports| ext_time
    eval_smoke_py -.->|imports| ext_torch
    eval_smoke_py -.->|imports| ext_topogpt3
    eval_temp_sweep_py -.->|imports| ext___future__
    eval_temp_sweep_py -.->|imports| ext_argparse
    eval_temp_sweep_py -.->|imports| ext_json
    eval_temp_sweep_py -.->|imports| ext_math
    eval_temp_sweep_py -.->|imports| ext_sys
    eval_temp_sweep_py -.->|imports| ext_time
    eval_temp_sweep_py -.->|imports| ext_pathlib
    eval_temp_sweep_py -.->|imports| ext_typing
    eval_temp_sweep_py -.->|imports| ext_torch
    ext_eval_noise_sweep["eval.noise_sweep"]
    class ext_eval_noise_sweep ext;
    eval_temp_sweep_py -.->|imports| ext_eval_noise_sweep
    eval_temp_sweep_py -.->|imports| ext_eval_harness
    infer_exploitgym_py -.->|imports| ext_argparse
    infer_exploitgym_py -.->|imports| ext_json
    ext_logging["logging"]
    class ext_logging ext;
    infer_exploitgym_py -.->|imports| ext_logging
    infer_exploitgym_py -.->|imports| ext_os
    infer_exploitgym_py -.->|imports| ext_sys
    infer_exploitgym_py -.->|imports| ext_pathlib
    infer_exploitgym_py -.->|imports| ext_torch
    infer_exploitgym_py -.->|imports| ext_torch_nn_functional
    infer_exploitgym_py -.->|imports| ext_safetensors_torch
    infer_exploitgym_py -.->|imports| ext_topogpt3_model
    infer_exploitgym_py -.->|imports| ext_topogpt3_train
    synthetic_dataset_py -.->|imports| ext_os
    synthetic_dataset_py -.->|imports| ext_sys
    synthetic_dataset_py -.->|imports| ext_json
    synthetic_dataset_py -.->|imports| ext_time
    ext_hashlib["hashlib"]
    class ext_hashlib ext;
    synthetic_dataset_py -.->|imports| ext_hashlib
    synthetic_dataset_py -.->|imports| ext_logging
    synthetic_dataset_py -.->|imports| ext_argparse
    synthetic_dataset_py -.->|imports| ext_tempfile
    synthetic_dataset_py -.->|imports| ext_pathlib
    synthetic_dataset_py -.->|imports| ext_typing
    synthetic_dataset_py -.->|imports| ext_dataclasses
    ext_datetime["datetime"]
    class ext_datetime ext;
    synthetic_dataset_py -.->|imports| ext_datetime
    synthetic_dataset_py -.->|imports| ext_threading
    ext_queue["queue"]
    class ext_queue ext;
    synthetic_dataset_py -.->|imports| ext_queue
    ext_concurrent_futures["concurrent.futures"]
    class ext_concurrent_futures ext;
    synthetic_dataset_py -.->|imports| ext_concurrent_futures
    synthetic_dataset_py -.->|imports| ext_torch
    synthetic_dataset_py -.->|imports| ext_numpy
    ext_tiktoken["tiktoken"]
    class ext_tiktoken ext;
    synthetic_dataset_py -.->|imports| ext_tiktoken
    ext_requests["requests"]
    class ext_requests ext;
    synthetic_dataset_py -.->|imports| ext_requests
    synthetic_dataset_py -.->|imports| ext_requests
    synthetic_dataset_py -.->|imports| ext_requests
    synthetic_dataset_py -.->|imports| ext_requests
    synthetic_dataset_py -.->|imports| ext_requests
    synthetic_dataset_py -.->|imports| ext_requests
    tests_test_jlens_py -.->|imports| ext___future__
    ext_pytest["pytest"]
    class ext_pytest ext;
    tests_test_jlens_py -.->|imports| ext_pytest
    tests_test_jlens_py -.->|imports| ext_torch
    ext_topogpt3_lens_model["topogpt3.lens_model"]
    class ext_topogpt3_lens_model ext;
    tests_test_jlens_py -.->|imports| ext_topogpt3_lens_model
    ext_topogpt3_jlens["topogpt3.jlens"]
    class ext_topogpt3_jlens ext;
    tests_test_jlens_py -.->|imports| ext_topogpt3_jlens
    tests_test_lens_model_py -.->|imports| ext___future__
    tests_test_lens_model_py -.->|imports| ext_pytest
    tests_test_lens_model_py -.->|imports| ext_torch
    tests_test_lens_model_py -.->|imports| ext_topogpt3_lens_model
    tests_test_lens_model_py -.->|imports| ext_topogpt3_model
    tests_test_lens_model_py -.->|imports| ext_topogpt3_model
    ext_types["types"]
    class ext_types ext;
    tests_test_lens_model_py -.->|imports| ext_types
    tests_test_lens_model_py -.->|imports| ext_topogpt3_jlens
    tests_test_lens_model_py -.->|imports| ext_topogpt3_jlens
    tests_test_lens_model_py -.->|imports| ext_topogpt3_jlens
    tests_test_lens_model_py -.->|imports| ext_topogpt3_jlens
    topogpt3___init___py -.->|imports| ext___future__
    ext_model["model"]
    class ext_model ext;
    topogpt3___init___py -.->|imports| ext_model
    ext_train["train"]
    class ext_train ext;
    topogpt3___init___py -.->|imports| ext_train
    ext_merged_config["merged_config"]
    class ext_merged_config ext;
    topogpt3___init___py -.->|imports| ext_merged_config
    ext_inference["inference"]
    class ext_inference ext;
    topogpt3___init___py -.->|imports| ext_inference
    ext_inference_hrm["inference_hrm"]
    class ext_inference_hrm ext;
    topogpt3___init___py -.->|imports| ext_inference_hrm
    ext_lens_model["lens_model"]
    class ext_lens_model ext;
    topogpt3___init___py -.->|imports| ext_lens_model
    ext_jlens["jlens"]
    class ext_jlens ext;
    topogpt3___init___py -.->|imports| ext_jlens
    topogpt3___main___py -.->|imports| ext___future__
    topogpt3___main___py -.->|imports| ext_sys
    topogpt3___main___py -.->|imports| ext_jlens
    topogpt3___main___py -.->|imports| ext_lens_model
    ext_api_server["api_server"]
    class ext_api_server ext;
    topogpt3___main___py -.->|imports| ext_api_server
    topogpt3___main___py -.->|imports| ext_inference
    topogpt3___main___py -.->|imports| ext_inference_hrm
    topogpt3___main___py -.->|imports| ext_train
    topogpt3_api_server_py -.->|imports| ext___future__
    topogpt3_api_server_py -.->|imports| ext_argparse
    topogpt3_api_server_py -.->|imports| ext_hashlib
    ext_hmac["hmac"]
    class ext_hmac ext;
    topogpt3_api_server_py -.->|imports| ext_hmac
    topogpt3_api_server_py -.->|imports| ext_json
    topogpt3_api_server_py -.->|imports| ext_logging
    topogpt3_api_server_py -.->|imports| ext_os
    topogpt3_api_server_py -.->|imports| ext_re
    ext_secrets["secrets"]
    class ext_secrets ext;
    topogpt3_api_server_py -.->|imports| ext_secrets
    topogpt3_api_server_py -.->|imports| ext_sys
    topogpt3_api_server_py -.->|imports| ext_time
    topogpt3_api_server_py -.->|imports| ext_collections
    topogpt3_api_server_py -.->|imports| ext_contextlib
    topogpt3_api_server_py -.->|imports| ext_dataclasses
    topogpt3_api_server_py -.->|imports| ext_pathlib
    topogpt3_api_server_py -.->|imports| ext_typing
    topogpt3_api_server_py -.->|imports| ext_torch
    topogpt3_api_server_py -.->|imports| ext_model
    topogpt3_api_server_py -.->|imports| ext_safetensors_torch
    ext_fastapi["fastapi"]
    class ext_fastapi ext;
    topogpt3_api_server_py -.->|imports| ext_fastapi
    ext_fastapi_middleware_cors["fastapi.middleware.cors"]
    class ext_fastapi_middleware_cors ext;
    topogpt3_api_server_py -.->|imports| ext_fastapi_middleware_cors
    ext_fastapi_middleware_gzip["fastapi.middleware.gzip"]
    class ext_fastapi_middleware_gzip ext;
    topogpt3_api_server_py -.->|imports| ext_fastapi_middleware_gzip
    ext_fastapi_responses["fastapi.responses"]
    class ext_fastapi_responses ext;
    topogpt3_api_server_py -.->|imports| ext_fastapi_responses
    ext_pydantic["pydantic"]
    class ext_pydantic ext;
    topogpt3_api_server_py -.->|imports| ext_pydantic
    ext_uvicorn["uvicorn"]
    class ext_uvicorn ext;
    topogpt3_api_server_py -.->|imports| ext_uvicorn
    topogpt3_api_server_py -.->|imports| ext_safetensors
    ext_continuation["continuation"]
    class ext_continuation ext;
    topogpt3_api_server_py -.->|imports| ext_continuation
    topogpt3_continuation_py -.->|imports| ext___future__
    topogpt3_continuation_py -.->|imports| ext_re
    topogpt3_continuation_py -.->|imports| ext_typing
    topogpt3_ewc_py -.->|imports| ext___future__
    topogpt3_ewc_py -.->|imports| ext_logging
    topogpt3_ewc_py -.->|imports| ext_typing
    topogpt3_ewc_py -.->|imports| ext_torch
    ext_torch_nn["torch.nn"]
    class ext_torch_nn ext;
    topogpt3_ewc_py -.->|imports| ext_torch_nn
    topogpt3_ewc_py -.->|imports| ext_torch_nn_functional
    topogpt3_exploitgym_config_py -.->|imports| ext___future__
    topogpt3_exploitgym_config_py -.->|imports| ext_dataclasses
    topogpt3_exploitgym_config_py -.->|imports| ext_typing
    topogpt3_exploitgym_config_py -.->|imports| ext_model
    ext_exploitgym_loader["exploitgym_loader"]
    class ext_exploitgym_loader ext;
    topogpt3_exploitgym_config_py -.->|imports| ext_exploitgym_loader
    topogpt3_exploitgym_loader_py -.->|imports| ext___future__
    topogpt3_exploitgym_loader_py -.->|imports| ext_json
    topogpt3_exploitgym_loader_py -.->|imports| ext_logging
    topogpt3_exploitgym_loader_py -.->|imports| ext_os
    topogpt3_exploitgym_loader_py -.->|imports| ext_subprocess
    topogpt3_exploitgym_loader_py -.->|imports| ext_time
    topogpt3_exploitgym_loader_py -.->|imports| ext_dataclasses
    topogpt3_exploitgym_loader_py -.->|imports| ext_pathlib
    topogpt3_exploitgym_loader_py -.->|imports| ext_typing
    topogpt3_exploitgym_loader_py -.->|imports| ext_numpy
    topogpt3_exploitgym_loader_py -.->|imports| ext_model
    topogpt3_hodge_cm_py -.->|imports| ext_math
    topogpt3_hodge_cm_py -.->|imports| ext_torch
    topogpt3_hodge_cm_py -.->|imports| ext_torch_nn
    topogpt3_hodge_cm_py -.->|imports| ext_torch_nn_functional
    topogpt3_inference_py -.->|imports| ext___future__
    topogpt3_inference_py -.->|imports| ext_argparse
    topogpt3_inference_py -.->|imports| ext_logging
    topogpt3_inference_py -.->|imports| ext_sys
    topogpt3_inference_py -.->|imports| ext_time
    topogpt3_inference_py -.->|imports| ext_dataclasses
    topogpt3_inference_py -.->|imports| ext_pathlib
    topogpt3_inference_py -.->|imports| ext_typing
    topogpt3_inference_py -.->|imports| ext_torch
    topogpt3_inference_py -.->|imports| ext_safetensors
    topogpt3_inference_py -.->|imports| ext_safetensors_torch
    topogpt3_inference_hrm_py -.->|imports| ext___future__
    topogpt3_inference_hrm_py -.->|imports| ext_argparse
    topogpt3_inference_hrm_py -.->|imports| ext_logging
    topogpt3_inference_hrm_py -.->|imports| ext_sys
    topogpt3_inference_hrm_py -.->|imports| ext_time
    topogpt3_inference_hrm_py -.->|imports| ext_dataclasses
    topogpt3_inference_hrm_py -.->|imports| ext_pathlib
    topogpt3_inference_hrm_py -.->|imports| ext_typing
    topogpt3_inference_hrm_py -.->|imports| ext_torch
    topogpt3_inference_hrm_py -.->|imports| ext_torch_nn_functional
    topogpt3_inference_hrm_py -.->|imports| ext_safetensors
    topogpt3_inference_hrm_py -.->|imports| ext_safetensors_torch
    topogpt3_inference_hrm_py -.->|imports| ext_continuation
    topogpt3_jlens_py -.->|imports| ext___future__
    topogpt3_jlens_py -.->|imports| ext_logging
    topogpt3_jlens_py -.->|imports| ext_math
    topogpt3_jlens_py -.->|imports| ext_os
    topogpt3_jlens_py -.->|imports| ext_time
    ext_collections_abc["collections.abc"]
    class ext_collections_abc ext;
    topogpt3_jlens_py -.->|imports| ext_collections_abc
    topogpt3_jlens_py -.->|imports| ext_dataclasses
    topogpt3_jlens_py -.->|imports| ext_typing
    topogpt3_jlens_py -.->|imports| ext_torch
    topogpt3_jlens_py -.->|imports| ext_torch
    topogpt3_jlens_py -.->|imports| ext_lens_model
    topogpt3_jlens_py -.->|imports| ext_argparse
    topogpt3_jlens_py -.->|imports| ext_lens_model
    ext_huggingface_hub["huggingface_hub"]
    class ext_huggingface_hub ext;
    topogpt3_jlens_py -.->|imports| ext_huggingface_hub
    topogpt3_lens_model_py -.->|imports| ext___future__
    topogpt3_lens_model_py -.->|imports| ext_json
    topogpt3_lens_model_py -.->|imports| ext_collections_abc
    topogpt3_lens_model_py -.->|imports| ext_dataclasses
    topogpt3_lens_model_py -.->|imports| ext_pathlib
    topogpt3_lens_model_py -.->|imports| ext_types
    topogpt3_lens_model_py -.->|imports| ext_typing
    topogpt3_lens_model_py -.->|imports| ext_torch
    topogpt3_lens_model_py -.->|imports| ext_torch
    topogpt3_lens_model_py -.->|imports| ext_safetensors_torch
    topogpt3_lens_model_py -.->|imports| ext_time
    topogpt3_lens_model_py -.->|imports| ext_torch
    topogpt3_lens_model_py -.->|imports| ext_model
    topogpt3_lens_model_py -.->|imports| ext_model
    topogpt3_merged_config_py -.->|imports| ext___future__
    topogpt3_merged_config_py -.->|imports| ext_dataclasses
    topogpt3_merged_config_py -.->|imports| ext_typing
    topogpt3_merged_config_py -.->|imports| ext_model
    topogpt3_merged_config_py -.->|imports| ext_exploitgym_loader
    topogpt3_model_py -.->|imports| ext_torch
    topogpt3_model_py -.->|imports| ext_torch_nn
    topogpt3_model_py -.->|imports| ext_torch_nn_functional
    ext_torch_utils_checkpoint["torch.utils.checkpoint"]
    class ext_torch_utils_checkpoint ext;
    topogpt3_model_py -.->|imports| ext_torch_utils_checkpoint
    topogpt3_model_py -.->|imports| ext_safetensors_torch
    topogpt3_model_py -.->|imports| ext_numpy
    topogpt3_model_py -.->|imports| ext_math
    topogpt3_model_py -.->|imports| ext_os
    topogpt3_model_py -.->|imports| ext_sys
    topogpt3_model_py -.->|imports| ext_time
    topogpt3_model_py -.->|imports| ext_pathlib
    topogpt3_model_py -.->|imports| ext_json
    topogpt3_model_py -.->|imports| ext_hashlib
    topogpt3_model_py -.->|imports| ext_logging
    ext_warnings["warnings"]
    class ext_warnings ext;
    topogpt3_model_py -.->|imports| ext_warnings
    topogpt3_model_py -.->|imports| ext_argparse
    topogpt3_model_py -.->|imports| ext_datetime
    topogpt3_model_py -.->|imports| ext_typing
    topogpt3_model_py -.->|imports| ext_dataclasses
    topogpt3_model_py -.->|imports| ext_collections
    topogpt3_model_py -.->|imports| ext_collections
    ext_synthetic_dataset["synthetic_dataset"]
    class ext_synthetic_dataset ext;
    topogpt3_model_py -.->|imports| ext_synthetic_dataset
    topogpt3_model_py -.->|imports| ext_continuation
    topogpt3_model_py -.->|imports| ext_tiktoken
    topogpt3_model_py -.->|imports| ext_shutil
    topogpt3_train_py -.->|imports| ext___future__
    topogpt3_train_py -.->|imports| ext_argparse
    topogpt3_train_py -.->|imports| ext_json
    topogpt3_train_py -.->|imports| ext_logging
    topogpt3_train_py -.->|imports| ext_math
    topogpt3_train_py -.->|imports| ext_os
    topogpt3_train_py -.->|imports| ext_sys
    topogpt3_train_py -.->|imports| ext_time
    topogpt3_train_py -.->|imports| ext_collections
    topogpt3_train_py -.->|imports| ext_dataclasses
    topogpt3_train_py -.->|imports| ext_datetime
    topogpt3_train_py -.->|imports| ext_pathlib
    topogpt3_train_py -.->|imports| ext_typing
    topogpt3_train_py -.->|imports| ext_numpy
    topogpt3_train_py -.->|imports| ext_torch
    topogpt3_train_py -.->|imports| ext_torch_nn
    topogpt3_train_py -.->|imports| ext_torch_nn_functional
    topogpt3_train_py -.->|imports| ext_model
    topogpt3_train_py -.->|imports| ext_exploitgym_loader
    ext_exploitgym_config["exploitgym_config"]
    class ext_exploitgym_config ext;
    topogpt3_train_py -.->|imports| ext_exploitgym_config
    topogpt3_train_py -.->|imports| ext_merged_config
    ext_ewc["ewc"]
    class ext_ewc ext;
    topogpt3_train_py -.->|imports| ext_ewc
    topogpt3_train_py -.->|imports| ext_safetensors_torch
    topogpt3_train_py -.->|imports| ext_safetensors_torch
    topogpt3_train_py -.->|imports| ext_datasets
    topogpt3_train_py -.->|imports| ext_safetensors_torch
    topogpt3_vanilla_control_py -.->|imports| ext_argparse
    topogpt3_vanilla_control_py -.->|imports| ext_math
    topogpt3_vanilla_control_py -.->|imports| ext_pathlib
    topogpt3_vanilla_control_py -.->|imports| ext_torch
    topogpt3_vanilla_control_py -.->|imports| ext_torch_nn
    topogpt3_vanilla_control_py -.->|imports| ext_torch_nn_functional
    topogpt3_vanilla_control_py -.->|imports| ext_dataclasses
    topogpt3_vanilla_control_py -.->|imports| ext_torch
    topogpt3_vanilla_control_py -.->|imports| ext_numpy
    transfer_weights_py -.->|imports| ext_argparse
    transfer_weights_py -.->|imports| ext_sys
    transfer_weights_py -.->|imports| ext_pathlib
    transfer_weights_py -.->|imports| ext_torch
    transfer_weights_py -.->|imports| ext_numpy
    transfer_weights_py -.->|imports| ext_topogpt3_model
    transfer_weights_py -.->|imports| ext_safetensors_torch
    transfer_weights_py -.->|imports| ext_safetensors_torch
```

---

## UML Class Diagram

Auto-generated Mermaid class diagram from parsed class-level symbols. Shows classes, structs, interfaces, traits, and their methods with inheritance and dependency relationships.

```mermaid
classDiagram
  class governor_py_TokenStream {
    <<class>>
    +make_loop_detector(window, min_repeats)
    +make_timeout_hook(per_token_s)
    +__init__(self)
    +put(self, tok)
    +mark_done(self)
    +drain(self)
    +wait_for_new(self, timeout)
    +is_closed(self)
    +__len__(self)
    +__post_init__(self)
  }
  class governor_py_StopReason {
    <<class>>
    +make_loop_detector(window, min_repeats)
    +make_timeout_hook(per_token_s)
    +__init__(self)
    +put(self, tok)
    +mark_done(self)
    +drain(self)
    +wait_for_new(self, timeout)
    +is_closed(self)
    +__len__(self)
    +__post_init__(self)
  }
  class governor_py_GenerationResult {
    <<class>>
    +make_loop_detector(window, min_repeats)
    +make_timeout_hook(per_token_s)
    +__init__(self)
    +put(self, tok)
    +mark_done(self)
    +drain(self)
    +wait_for_new(self, timeout)
    +is_closed(self)
    +__len__(self)
    +__post_init__(self)
  }
  class governor_py_GenerationGovernor {
    <<class>>
    +make_loop_detector(window, min_repeats)
    +make_timeout_hook(per_token_s)
    +__init__(self)
    +put(self, tok)
    +mark_done(self)
    +drain(self)
    +wait_for_new(self, timeout)
    +is_closed(self)
    +__len__(self)
    +__post_init__(self)
  }
  class harness_py_ModelLoader {
    <<class>>
    +load_humaneval(cache_dir)
    +build_prompt(problem)
    +extract_candidate(prompt, completion)
    +run_one_test(problem, candidate_src, timeout)
    +run_one_test_sandboxed(problem, candidate_src, timeout, sandbox_cfg)
    +make_sampler(mode, settings_kwargs)
    +completion_for_problem(sampler, prompt)
    +evaluate_problem(problem, loader, args, sample_idx)
    +main()
    +__init__(self, ckpt_dir, ckpt_name, device)
  }
  class sandbox_py_SandboxConfig {
    <<class>>
    +_names_imported(tree)
    +_blocked_dunder_access(tree, blocked)
    +_max_depth(tree)
    +check_safety(source, cfg)
    +_build_worker_src(allowed_builtin_names, program_src, blocked_modules)
    +safe_exec(program_src, cfg, extra_globals)
    +describe_policy(cfg)
    +d(node, cur)
  }
  class synthetic_dataset_py_LLMBackend {
    <<class>>
    +build_backend(provider, model)
    +validate_sample(sample)
    +build_logger(level)
    +parse_args()
    +load_paths(paths_arg, paths_file, max_files)
    +main()
    +generate(self, prompt)
    +name(self)
    +__init__(self, model, api_key, max_tokens, temperature, timeout)
    +name(self)
  }
  class synthetic_dataset_py_GroqBackend {
    <<class>>
    +build_backend(provider, model)
    +validate_sample(sample)
    +build_logger(level)
    +parse_args()
    +load_paths(paths_arg, paths_file, max_files)
    +main()
    +generate(self, prompt)
    +name(self)
    +__init__(self, model, api_key, max_tokens, temperature, timeout)
    +name(self)
  }
  class synthetic_dataset_py_OpenRouterBackend {
    <<class>>
    +build_backend(provider, model)
    +validate_sample(sample)
    +build_logger(level)
    +parse_args()
    +load_paths(paths_arg, paths_file, max_files)
    +main()
    +generate(self, prompt)
    +name(self)
    +__init__(self, model, api_key, max_tokens, temperature, timeout)
    +name(self)
  }
  class synthetic_dataset_py_OllamaBackend {
    <<class>>
    +build_backend(provider, model)
    +validate_sample(sample)
    +build_logger(level)
    +parse_args()
    +load_paths(paths_arg, paths_file, max_files)
    +main()
    +generate(self, prompt)
    +name(self)
    +__init__(self, model, api_key, max_tokens, temperature, timeout)
    +name(self)
  }
  class synthetic_dataset_py_ProcessedManifest {
    <<class>>
    +build_backend(provider, model)
    +validate_sample(sample)
    +build_logger(level)
    +parse_args()
    +load_paths(paths_arg, paths_file, max_files)
    +main()
    +generate(self, prompt)
    +name(self)
    +__init__(self, model, api_key, max_tokens, temperature, timeout)
    +name(self)
  }
  class synthetic_dataset_py_SyntheticDatasetGenerator {
    <<class>>
    +build_backend(provider, model)
    +validate_sample(sample)
    +build_logger(level)
    +parse_args()
    +load_paths(paths_arg, paths_file, max_files)
    +main()
    +generate(self, prompt)
    +name(self)
    +__init__(self, model, api_key, max_tokens, temperature, timeout)
    +name(self)
  }
  class test_jlens_py_TestValidPositionMask {
    <<class>>
    +test_basic_mask(self)
    +test_too_short_raises(self)
    +test_negative_skip_raises(self)
    +test_all_positions_valid(self)
    +test_exact_minimum_length(self)
    +model(self)
    +test_returns_jacobians_for_source_layers(self, model)
    +test_late_layer_jacobian_close_to_identity(self, model)
    +test_earlier_layers_further_from_identity(self, model)
    +test_exact_jacobian_for_last_block(self, model)
  }
  class test_jlens_py_TestJacobianForPrompt {
    <<class>>
    +test_basic_mask(self)
    +test_too_short_raises(self)
    +test_negative_skip_raises(self)
    +test_all_positions_valid(self)
    +test_exact_minimum_length(self)
    +model(self)
    +test_returns_jacobians_for_source_layers(self, model)
    +test_late_layer_jacobian_close_to_identity(self, model)
    +test_earlier_layers_further_from_identity(self, model)
    +test_exact_jacobian_for_last_block(self, model)
  }
  class test_jlens_py_TestFit {
    <<class>>
    +test_basic_mask(self)
    +test_too_short_raises(self)
    +test_negative_skip_raises(self)
    +test_all_positions_valid(self)
    +test_exact_minimum_length(self)
    +model(self)
    +test_returns_jacobians_for_source_layers(self, model)
    +test_late_layer_jacobian_close_to_identity(self, model)
    +test_earlier_layers_further_from_identity(self, model)
    +test_exact_jacobian_for_last_block(self, model)
  }
  class test_jlens_py_TestJacobianLens {
    <<class>>
    +test_basic_mask(self)
    +test_too_short_raises(self)
    +test_negative_skip_raises(self)
    +test_all_positions_valid(self)
    +test_exact_minimum_length(self)
    +model(self)
    +test_returns_jacobians_for_source_layers(self, model)
    +test_late_layer_jacobian_close_to_identity(self, model)
    +test_earlier_layers_further_from_identity(self, model)
    +test_exact_jacobian_for_last_block(self, model)
  }
  class test_jlens_py_TestFitCheckpoint {
    <<class>>
    +test_basic_mask(self)
    +test_too_short_raises(self)
    +test_negative_skip_raises(self)
    +test_all_positions_valid(self)
    +test_exact_minimum_length(self)
    +model(self)
    +test_returns_jacobians_for_source_layers(self, model)
    +test_late_layer_jacobian_close_to_identity(self, model)
    +test_earlier_layers_further_from_identity(self, model)
    +test_exact_jacobian_for_last_block(self, model)
  }
  class test_jlens_py_TestConfig {
    <<class>>
    +test_basic_mask(self)
    +test_too_short_raises(self)
    +test_negative_skip_raises(self)
    +test_all_positions_valid(self)
    +test_exact_minimum_length(self)
    +model(self)
    +test_returns_jacobians_for_source_layers(self, model)
    +test_late_layer_jacobian_close_to_identity(self, model)
    +test_earlier_layers_further_from_identity(self, model)
    +test_exact_jacobian_for_last_block(self, model)
  }
  class test_jlens_py_TestTopoGPT3JLensAppConfig {
    <<class>>
    +test_basic_mask(self)
    +test_too_short_raises(self)
    +test_negative_skip_raises(self)
    +test_all_positions_valid(self)
    +test_exact_minimum_length(self)
    +model(self)
    +test_returns_jacobians_for_source_layers(self, model)
    +test_late_layer_jacobian_close_to_identity(self, model)
    +test_earlier_layers_further_from_identity(self, model)
    +test_exact_jacobian_for_last_block(self, model)
  }
  class test_lens_model_py_TestTopoGPT3LensConfig {
    <<class>>
    +test_default_config(self)
    +test_from_topogpt2_config(self)
    +test_probe_checkpoint_missing_raises(self, tmp_path)
    +test_default_parameters(self)
    +test_forward_output_shape(self)
    +test_weight_tied(self)
    +raw_model(self)
    +lens_model(self, raw_model)
    +test_exposes_protocol_attributes(self, lens_model, raw_model)
    +test_encode_text_to_token_ids(self, lens_model)
  }
  class test_lens_model_py_TestTinyDecoder {
    <<class>>
    +test_default_config(self)
    +test_from_topogpt2_config(self)
    +test_probe_checkpoint_missing_raises(self, tmp_path)
    +test_default_parameters(self)
    +test_forward_output_shape(self)
    +test_weight_tied(self)
    +raw_model(self)
    +lens_model(self, raw_model)
    +test_exposes_protocol_attributes(self, lens_model, raw_model)
    +test_encode_text_to_token_ids(self, lens_model)
  }
  class test_lens_model_py_TestTopoGPT3LensModel {
    <<class>>
    +test_default_config(self)
    +test_from_topogpt2_config(self)
    +test_probe_checkpoint_missing_raises(self, tmp_path)
    +test_default_parameters(self)
    +test_forward_output_shape(self)
    +test_weight_tied(self)
    +raw_model(self)
    +lens_model(self, raw_model)
    +test_exposes_protocol_attributes(self, lens_model, raw_model)
    +test_encode_text_to_token_ids(self, lens_model)
  }
  class test_lens_model_py_TestTopoGPT3LensModelWithRecording {
    <<class>>
    +test_default_config(self)
    +test_from_topogpt2_config(self)
    +test_probe_checkpoint_missing_raises(self, tmp_path)
    +test_default_parameters(self)
    +test_forward_output_shape(self)
    +test_weight_tied(self)
    +raw_model(self)
    +lens_model(self, raw_model)
    +test_exposes_protocol_attributes(self, lens_model, raw_model)
    +test_encode_text_to_token_ids(self, lens_model)
  }
  class test_lens_model_py_TestTopoGPT3LensModelEdgeCases {
    <<class>>
    +test_default_config(self)
    +test_from_topogpt2_config(self)
    +test_probe_checkpoint_missing_raises(self, tmp_path)
    +test_default_parameters(self)
    +test_forward_output_shape(self)
    +test_weight_tied(self)
    +raw_model(self)
    +lens_model(self, raw_model)
    +test_exposes_protocol_attributes(self, lens_model, raw_model)
    +test_encode_text_to_token_ids(self, lens_model)
  }
  class api_server_py_ApiKey {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class api_server_py_AuthState {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class api_server_py_TokenBucket {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class api_server_py_RateLimiter {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class api_server_py_IpBanner {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class api_server_py_CompletionRequest {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class api_server_py_Message {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class api_server_py_ChatCompletionRequest {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class api_server_py_ServerModel {
    <<class>>
    +_setup_logging(verbose)
    +_parse_keys(raw)
    +_sha256(raw)
    +_sanitize_stop(stop)
    +_resolve_device(device)
    +_probe_arch(checkpoint_dir)
    +load_model(checkpoint, device)
    +lifespan(app)
    +_security_middleware(request, call_next)
    +_real_ip(request)
  }
  class ewc_py_EWCRegularizer {
    <<class>>
    +__init__(self, model, lambda_ewc, fisher_num_samples)
    +compute_fisher(self, dataloader, vocab_size, device, max_batches)
    +penalty(self, model)
    +save_state(self, path)
    +load_state(self, path)
    +__init__(self, max_tiers, replay_ratio)
    +add_tier(self, tier_index, dataloader)
    +_get_next_batch(self, tier_index, dataloader)
    +sample_batch(self, current_batch_size)
  }
  class ewc_py_ReplayBuffer {
    <<class>>
    +__init__(self, model, lambda_ewc, fisher_num_samples)
    +compute_fisher(self, dataloader, vocab_size, device, max_batches)
    +penalty(self, model)
    +save_state(self, path)
    +load_state(self, path)
    +__init__(self, max_tiers, replay_ratio)
    +add_tier(self, tier_index, dataloader)
    +_get_next_batch(self, tier_index, dataloader)
    +sample_batch(self, current_batch_size)
  }
  class exploitgym_config_py_TopoExploitConfig {
    <<class>>
    +build_topogpt2_config(self, max_seq_len, attn_window)
    +build_loader_config(self)
  }
  class exploitgym_loader_py_ExploitGymLoaderConfig {
    <<class>>
    +ensure_repo(repo_cache, logger)
    +_read_file_safe(path, max_chars)
    +_parse_task_ids(repo_path)
    +_task_id_to_path(task_id)
    +parse_task(task_dir, task_id, max_chars)
    +_format_vulnerability_analysis(task)
    +_format_patch_analysis(task)
    +_format_exploit_development(task)
    +__init__(self, config, tokenizer, logger)
    +_tier_paths(self, tier)
  }
  class exploitgym_loader_py_ExploitGymDataLoader {
    <<class>>
    +ensure_repo(repo_cache, logger)
    +_read_file_safe(path, max_chars)
    +_parse_task_ids(repo_path)
    +_task_id_to_path(task_id)
    +parse_task(task_dir, task_id, max_chars)
    +_format_vulnerability_analysis(task)
    +_format_patch_analysis(task)
    +_format_exploit_development(task)
    +__init__(self, config, tokenizer, logger)
    +_tier_paths(self, tier)
  }
  class hodge_cm_py_CMPhaseQuantSTE {
    <<class>>
    +cm_phase_loss(kr, ki, m)
    +apply_cm_soft(kr, ki, m, beta)
    +freq_grid_laplacian(fh, fw, device, dtype)
    +hodge_heat_residue(kr, ki, t)
    +harmonic_ratio(kr, ki, t)
    +fisher_hinge(singulars, r, margin)
    +langevin_refine(param, energy_fn, steps, eta, temp, seed)
    +gibbs_cluster_assign(feats, n_clusters, iters, seed)
    +forward(ctx, phase, m, beta)
    +backward(ctx, grad_out)
  }
  class hodge_cm_py_HeckeScalePool {
    <<class>>
    +cm_phase_loss(kr, ki, m)
    +apply_cm_soft(kr, ki, m, beta)
    +freq_grid_laplacian(fh, fw, device, dtype)
    +hodge_heat_residue(kr, ki, t)
    +harmonic_ratio(kr, ki, t)
    +fisher_hinge(singulars, r, margin)
    +langevin_refine(param, energy_fn, steps, eta, temp, seed)
    +gibbs_cluster_assign(feats, n_clusters, iters, seed)
    +forward(ctx, phase, m, beta)
    +backward(ctx, grad_out)
  }
  class inference_py_ScalePreset {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_InferenceSettings {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_InferenceLoggerFactory {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_SecurePathResolver {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_SourceModuleLoader {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_CheckpointPaths {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_WeightShapeProbe {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_TopoGPT2ConfigAligner {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_TokenizerFactory {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
  class inference_py_GaussPatchApplier {
    <<class>>
    +main(argv)
    +scale_presets()
    +preset(self)
    +validate(self)
    +build(settings)
    +resolve_under(root)
    +require_existing_file(path, expected_suffix)
    +__init__(self, settings, logger)
    +load(self)
    +__init__(self, settings)
  }
```

---

## Code Property Graph

Machine-readable Code Property Graph (CPG) in JSON-LD format. This block allows AI agents to parse the full structural graph without additional file reads. Compatible with GraphRAG pipelines.

```json
{"@context": "https://schema.org", "analysis": {"communities": [{"cohesion": 0.643, "id": 0, "label": "topogpt3: model", "size": 12}, {"cohesion": 0.636, "id": 1, "label": "eval: harness", "size": 7}, {"cohesion": 0.312, "id": 2, "label": "eval: governor", "size": 6}, {"cohesion": 0.417, "id": 3, "label": "topogpt3: inference_hrm", "size": 5}, {"cohesion": 0.455, "id": 4, "label": "topogpt3: jlens", "size": 4}], "god_nodes": [{"node_id": "topogpt3/model.py", "score": 48.7}, {"node_id": "topogpt3/__init__.py", "score": 30.0}, {"node_id": "topogpt3/train.py", "score": 24.4}, {"node_id": "topogpt3/lens_model.py", "score": 14.9}, {"node_id": "topogpt3/inference_hrm.py", "score": 13.6}, {"node_id": "eval/harness.py", "score": 13.2}, {"node_id": "topogpt3/jlens.py", "score": 12.9}, {"node_id": "topogpt3/__main__.py", "score": 12.1}, {"node_id": "topogpt3/api_server.py", "score": 10.6}, {"node_id": "topogpt3/exploitgym_loader.py", "score": 9.9}], "surprising_connections": [{"hops": 6, "source": "eval/sandbox_smoke.py", "target": "topogpt3/hodge_cm.py"}, {"hops": 5, "source": "eval/governor.py", "target": "eval/sandbox_smoke.py"}, {"hops": 5, "source": "eval/governor.py", "target": "topogpt3/hodge_cm.py"}, {"hops": 5, "source": "eval/hodge_cm_ablation.py", "target": "eval/sandbox_smoke.py"}, {"hops": 5, "source": "eval/integration_smoke.py", "target": "topogpt3/hodge_cm.py"}]}, "edges": [{"confidence": "EXTRACTED", "relation": "imports", "source": "app.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "app.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "app.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "app.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "app.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "app.py", "target": "topogpt3"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze.py", "target": "re"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze_results.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze_results.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze_results.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze_results.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/analyze_results.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "topogpt3"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "topogpt3.model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/diag_static.py", "target": "topogpt3.train"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor.py", "target": "threading"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor.py", "target": "enum"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor.py", "target": "torch.nn.functional"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "threading"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "eval.governor"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "topogpt3"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "safetensors"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/governor_smoke.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "contextlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "io"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "re"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "signal"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "subprocess"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "traceback"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "topogpt3"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "eval.samplers"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "datasets"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "eval.sandbox"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "eval.samplers"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/harness.py", "target": "safetensors"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "numpy"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "topogpt3.hodge_cm"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/hodge_cm_ablation.py", "target": "topogpt3.model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/integration_smoke.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/integration_smoke.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/integration_smoke.py", "target": "eval.harness"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "ast"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "re"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_analysis.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "topogpt3"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "topogpt3.model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "safetensors"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/noise_sweep.py", "target": "eval.harness"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "contextlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "io"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "re"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "traceback"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "safetensors"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "topogpt3"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/repair.py", "target": "datasets"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "re"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "shutil"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "statistics"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/report.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/samplers.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/samplers.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/samplers.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/samplers.py", "target": "topogpt3"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "ast"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "subprocess"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "tempfile"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "textwrap"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox_smoke.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox_smoke.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/sandbox_smoke.py", "target": "eval.sandbox"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/smoke.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/smoke.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/smoke.py", "target": "topogpt3"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "eval.noise_sweep"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "eval/temp_sweep.py", "target": "eval.harness"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "torch.nn.functional"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "topogpt3.model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "infer_exploitgym.py", "target": "topogpt3.train"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "hashlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "tempfile"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "datetime"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "threading"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "queue"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "concurrent.futures"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "numpy"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "tiktoken"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "requests"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "requests"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "requests"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "requests"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "requests"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "synthetic_dataset.py", "target": "requests"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_jlens.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_jlens.py", "target": "pytest"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_jlens.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_jlens.py", "target": "topogpt3.lens_model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_jlens.py", "target": "topogpt3.jlens"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "pytest"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "topogpt3.lens_model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "topogpt3.model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "topogpt3.model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "types"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "topogpt3.jlens"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "topogpt3.jlens"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "topogpt3.jlens"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "tests/test_lens_model.py", "target": "topogpt3.jlens"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__init__.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__init__.py", "target": "model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__init__.py", "target": "train"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__init__.py", "target": "merged_config"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__init__.py", "target": "inference"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__init__.py", "target": "inference_hrm"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__init__.py", "target": "lens_model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__init__.py", "target": "jlens"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__main__.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__main__.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__main__.py", "target": "jlens"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__main__.py", "target": "lens_model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__main__.py", "target": "api_server"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__main__.py", "target": "inference"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__main__.py", "target": "inference_hrm"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/__main__.py", "target": "train"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "hashlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "hmac"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "re"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "secrets"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "contextlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "fastapi"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "fastapi.middleware.cors"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "fastapi.middleware.gzip"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "fastapi.responses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "pydantic"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "uvicorn"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "safetensors"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/api_server.py", "target": "continuation"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/continuation.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/continuation.py", "target": "re"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/continuation.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/ewc.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/ewc.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/ewc.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/ewc.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/ewc.py", "target": "torch.nn"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/ewc.py", "target": "torch.nn.functional"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_config.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_config.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_config.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_config.py", "target": "model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_config.py", "target": "exploitgym_loader"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "subprocess"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "numpy"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/exploitgym_loader.py", "target": "model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/hodge_cm.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/hodge_cm.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/hodge_cm.py", "target": "torch.nn"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/hodge_cm.py", "target": "torch.nn.functional"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "safetensors"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "torch.nn.functional"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "safetensors"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/inference_hrm.py", "target": "continuation"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "collections.abc"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "lens_model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "lens_model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/jlens.py", "target": "huggingface_hub"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "collections.abc"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "types"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/lens_model.py", "target": "model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/merged_config.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/merged_config.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/merged_config.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/merged_config.py", "target": "model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/merged_config.py", "target": "exploitgym_loader"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "torch.nn"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "torch.nn.functional"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "torch.utils.checkpoint"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "numpy"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "hashlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "warnings"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "datetime"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "synthetic_dataset"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "continuation"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "tiktoken"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/model.py", "target": "shutil"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "__future__"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "json"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "logging"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "os"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "time"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "collections"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "datetime"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "typing"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "numpy"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "torch.nn"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "torch.nn.functional"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "exploitgym_loader"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "exploitgym_config"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "merged_config"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "ewc"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "datasets"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/train.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "math"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "torch.nn"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "torch.nn.functional"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "dataclasses"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "topogpt3/vanilla_control.py", "target": "numpy"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "transfer_weights.py", "target": "argparse"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "transfer_weights.py", "target": "sys"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "transfer_weights.py", "target": "pathlib"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "transfer_weights.py", "target": "torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "transfer_weights.py", "target": "numpy"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "transfer_weights.py", "target": "topogpt3.model"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "transfer_weights.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "imports", "source": "transfer_weights.py", "target": "safetensors.torch"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "app.py", "target": "topogpt3/__init__.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/diag_static.py", "target": "topogpt3/__init__.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/diag_static.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/diag_static.py", "target": "topogpt3/train.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/governor_smoke.py", "target": "eval/governor.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/governor_smoke.py", "target": "topogpt3/__init__.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/harness.py", "target": "topogpt3/__init__.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/harness.py", "target": "eval/samplers.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/harness.py", "target": "eval/sandbox.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/harness.py", "target": "eval/samplers.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/hodge_cm_ablation.py", "target": "topogpt3/hodge_cm.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/hodge_cm_ablation.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/integration_smoke.py", "target": "eval/harness.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/noise_sweep.py", "target": "topogpt3/__init__.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/noise_sweep.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/noise_sweep.py", "target": "eval/harness.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/repair.py", "target": "topogpt3/__init__.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/samplers.py", "target": "topogpt3/__init__.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/sandbox_smoke.py", "target": "eval/sandbox.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/smoke.py", "target": "topogpt3/__init__.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/temp_sweep.py", "target": "eval/noise_sweep.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "eval/temp_sweep.py", "target": "eval/harness.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "infer_exploitgym.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "infer_exploitgym.py", "target": "topogpt3/train.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_jlens.py", "target": "topogpt3/lens_model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_jlens.py", "target": "topogpt3/jlens.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_lens_model.py", "target": "topogpt3/lens_model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_lens_model.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_lens_model.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_lens_model.py", "target": "topogpt3/jlens.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_lens_model.py", "target": "topogpt3/jlens.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_lens_model.py", "target": "topogpt3/jlens.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "tests/test_lens_model.py", "target": "topogpt3/jlens.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__init__.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__init__.py", "target": "topogpt3/train.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__init__.py", "target": "topogpt3/merged_config.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__init__.py", "target": "topogpt3/inference.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__init__.py", "target": "topogpt3/inference_hrm.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__init__.py", "target": "topogpt3/lens_model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__init__.py", "target": "topogpt3/jlens.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__main__.py", "target": "topogpt3/jlens.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__main__.py", "target": "topogpt3/lens_model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__main__.py", "target": "topogpt3/api_server.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__main__.py", "target": "topogpt3/inference.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__main__.py", "target": "topogpt3/inference_hrm.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/__main__.py", "target": "topogpt3/train.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/api_server.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/api_server.py", "target": "topogpt3/continuation.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/exploitgym_config.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/exploitgym_config.py", "target": "topogpt3/exploitgym_loader.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/exploitgym_loader.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/inference_hrm.py", "target": "topogpt3/continuation.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/jlens.py", "target": "topogpt3/lens_model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/jlens.py", "target": "topogpt3/lens_model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/lens_model.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/lens_model.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/merged_config.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/merged_config.py", "target": "topogpt3/exploitgym_loader.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/model.py", "target": "synthetic_dataset.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/model.py", "target": "topogpt3/continuation.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/train.py", "target": "topogpt3/model.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/train.py", "target": "topogpt3/exploitgym_loader.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/train.py", "target": "topogpt3/exploitgym_config.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/train.py", "target": "topogpt3/merged_config.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "topogpt3/train.py", "target": "topogpt3/ewc.py"}, {"confidence": "EXTRACTED", "relation": "resolved_imports", "source": "transfer_weights.py", "target": "topogpt3/model.py"}], "generator": "readmenator", "metadata": {"edge_count": 6109, "file_count": 44, "language_count": 2, "symbol_count": 824}, "nodes": [{"doc": "Drop-in entry point that demonstrates how to use the topogpt3 package.  This file lives outside the package on purpose. Copy it (or its sections) into your own project after running ``pip install topogpt3``. Three usage patterns are shown:  1. ``run_inference`` calls the standard autoregressive sampler. 2. ``run_inference_hrm`` calls the hierarchical recursive reasoning sampler that reuses the same checkpoint with no extra trained parameters. 3. ``run_training`` launches the full curriculum trainer.  The script's main() exposes them through a tiny ``--mode`` CLI so the file is runnable as-is for a quick smoke test once a checkpoint exists.", "id": "app.py", "kind": "module", "label": "app.py", "language": "py", "sha256": "d456da403bd5058f", "symbol_count": 5, "symbols": [{"doc": "Run the standard sampler and return the generated completion text.", "kind": "function", "line": 46, "name": "run_inference", "signature": "def run_inference(prompt, checkpoint_dir, checkpoint_name, max_new_tokens, temperature, top_k, repetition_penalty, device)"}, {"doc": "Run the hierarchical recursive sampler and return the completion.", "kind": "function", "line": 71, "name": "run_inference_hrm", "signature": "def run_inference_hrm(prompt, checkpoint_dir, checkpoint_name, max_new_tokens, temperature, top_k, repetition_penalty, high_level_iters, low_level_iters, low_level_window, device)"}, {"doc": "Run the full TopoGPT3 curriculum trainer.", "kind": "function", "line": 105, "name": "run_training", "signature": "def run_training(scale, start_tier, device, prepare_data)"}, {"doc": "Build the top-level CLI for this entry point script.", "kind": "function", "line": 121, "name": "_build_parser", "signature": "def _build_parser()"}, {"doc": "Entry point invoked when the file is executed as a script.", "kind": "function", "line": 159, "name": "main", "signature": "def main(argv)"}]}, {"doc": "Aggregate HumanEval result JSONL files into a summary table.  Reads one or more .jsonl files produced by harness.py and computes: - pass@1, pass@k (using the unbiased estimator from the HumanEval paper when k > 1) - mean latency, mean generation length, tok/s - per-error classification", "id": "eval/analyze.py", "kind": "module", "label": "analyze.py", "language": "py", "sha256": "351b50620aa106b7", "symbol_count": 5, "symbols": [{"doc": "Unbiased estimator from the HumanEval paper.\n\npass@k = 1 - C(n-c, k) / C(n, k)   if n - c >= k else 1.0\nn = total samples, c = correct samples, k = target", "kind": "function", "line": 21, "name": "pass_at_k", "signature": "def pass_at_k(n, c, k)"}, {"doc": "Heuristic single-label error classifier.", "kind": "function", "line": 32, "name": "classify_error", "signature": "def classify_error(msg, candidate_src)"}, {"kind": "function", "line": 56, "name": "load_jsonl", "signature": "def load_jsonl(path)"}, {"kind": "function", "line": 61, "name": "summarize", "signature": "def summarize(paths)"}, {"kind": "function", "line": 103, "name": "main", "signature": "def main()"}]}, {"doc": "Analyze a HumanEval JSONL produced by harness.py.  For each failed problem the report shows: - the prompt fed to the model - the generated candidate after extraction - the hidden test that failed - the captured stdout/stderr and traceback  This makes it easy to see *how* and *why* a candidate failed without re-running the harness.  Usage: python eval/analyze_results.py eval/runs/run.jsonl python eval/analyze_results.py eval/runs/run.jsonl --summary python eval/analyze_results.py eval/runs/run.jsonl --task-id HumanEval/0", "id": "eval/analyze_results.py", "kind": "module", "label": "analyze_results.py", "language": "py", "sha256": "aed01199698beec6", "symbol_count": 4, "symbols": [{"kind": "function", "line": 26, "name": "load_records", "signature": "def load_records(path)"}, {"kind": "function", "line": 31, "name": "summarize", "signature": "def summarize(records)"}, {"kind": "function", "line": 44, "name": "show_failures", "signature": "def show_failures(records, task_id)"}, {"kind": "function", "line": 82, "name": "main", "signature": "def main()"}]}, {"doc": "Diagnostico estatico de un checkpoint TopoGPT3 congelado.  Calcula sobre los pesos espectrales congelados (sin reentrenar):  kappa_F   = sigma_max / sigma_min  del kernel espectral apilado (proxy del condition number de la Grassmanniana) delta     = max |theta - round(theta)|  sobre los arg det de overlaps (cuantifica cuanto se \"discretizan\" las fases complejas) W         = (1/2pi) sum arg det <U_n | U_{n+1}>  (winding acumulado sobre barridos en frecuencia — sin trayectoria temporal real, usamos un barrido sintetico sobre los modos FFT) r         = rango dominante por elbow de los valores singulares sigma_*   = valores singulares principales  NOTA IMPORTANTE: Este script NO reentrena. Trabaja unicamente con los kernels espectrales cuaternionicos ya aprendidos. La \"trayectoria\" W se define barriendo sobre los modos de frecuencia (no sobre pasos de entrenamiento), asi que W aqui mide coherencia de fase intra-modelo, no winding temporal. Esta distincion se reporta explicitamente en el JSONL de salida.  Salida: eval/runs/diag_static_<timestamp>.jsonl", "id": "eval/diag_static.py", "kind": "module", "label": "diag_static.py", "language": "py", "sha256": "80ab610c55205ed7", "symbol_count": 5, "symbols": [{"doc": "Muestrea n_samples overlaps aleatorios <u_i | u_j> sobre los vectores\nsingulares de K y mide cuanto se aleja su fase arg del reticulo 2*pi*Z.\n\ndelta = max |theta/2pi - round(theta/2pi)| sobre la muestra.\n\nTambien devuelve:\n  delta_mean, delta_median, frac_near_integer (|.| < 0.05)", "kind": "function", "line": 49, "name": "phase_discretization", "signature": "def phase_discretization(K, n_samples, seed)"}, {"doc": "Como el checkpoint es estatico, no hay trayectoria temporal.\nConstruimos una pseudo-trayectoria deslizando una ventana sobre\nlos modos de frecuencia (filas de K) y acumulando arg det del\noverlap entre ventanas consecutivas.\n\nW = (1/2pi) sum_n arg det <U_{n} | U_{n+1}>", "kind": "function", "line": 95, "name": "synthetic_winding", "signature": "def synthetic_winding(K, n_windows, window_size)"}, {"kind": "function", "line": 144, "name": "static_kappa", "signature": "def static_kappa(K)"}, {"kind": "function", "line": 171, "name": "context_length_diagnostic", "signature": "def context_length_diagnostic(model, tracker, device, lengths)"}, {"kind": "function", "line": 248, "name": "main", "signature": "def main()"}]}, {"doc": "Streaming + governance for autoregressive generation.  Two classes that fix two real problems with the existing `topogpt3.inference` pipeline:  - `TokenStream` — a thread-safe queue that captures raw token IDs as they are produced by the model. Enables post-hoc prefix agreement and exact-match metrics that need the *raw* token stream (the current harness only stores the post-extracted candidate text, losing that information).  - `GenerationGovernor` — wraps `model.generate` and exposes stop hooks: per-token timeout, loop detection (last K tokens repeat), and a user-callable cancel. Returns a `GenerationResult` with the stop reason so callers can distinguish \"ran out of tokens\" from \"hit the safety hook\" from \"user aborted\".  This is a Python port of the patterns in `claude-code-main/src/utils/stream.ts` (Stream<T> AsyncIterator wrapper) and `claude-code-main/src/query/stopHooks.ts` (AsyncGenerator with `preventContinuation`). The TypeScript originals are 76 and 473 lines respectively; this module is ~150 lines because Python's GIL lets us avoid the manual promise queueing.  NOTE: This module does NOT modify `topogpt3/model.py`. The generation loop is replicated here (not monkey-patched) so the original `generate` remains the single source of truth for the production sampler.", "id": "eval/governor.py", "kind": "module", "label": "governor.py", "language": "py", "sha256": "0a2e0b070f049594", "symbol_count": 20, "symbols": [{"doc": "Thread-safe single-producer / single-consumer queue of token IDs.\n\nThe producer (the generation loop) calls `put(tok)` for each new\ntoken. Consumers can iterate via `iter_tokens(block=True)` or\n`drain()` to get everything emitted so far.\n\nThe stream tracks a monotonic counter so consumers can detect\n\"no new tokens since last call\" cheaply.", "kind": "class", "line": 45, "name": "TokenStream", "signature": "class TokenStream"}, {"kind": "class", "line": 99, "name": "StopReason", "signature": "class StopReason(str, Enum)"}, {"doc": "Outcome of a governed generation.", "kind": "class", "line": 109, "name": "GenerationResult", "signature": "class GenerationResult"}, {"doc": "Run a model's autoregressive generation loop with optional stop\nhooks and a streaming interface.\n\nUsage:\n    ts = TokenStream()\n    governor = GenerationGovernor(\n        model=model,\n        ctx=prompt_tensor,\n        stream=ts,\n        max_new_tokens=256,\n        temperature=0.2,\n        top_k=40,\n        repetition_penalty=1.1,\n    )\n    result = governor.run(stop_hooks=[loop_detector, timeout_hook])\n    if result.stop_reason == StopReason.LOOP:\n        ...", "kind": "class", "line": 134, "name": "GenerationGovernor", "signature": "class GenerationGovernor"}, {"doc": "Return True if the last `window` tokens contain a sub-sequence\nof length >= `min_repeats` that repeats consecutively.\n\nCatches the \"model is stuck in a loop\" pathology where a 24M-param\nmodel emits the same 4-token pattern indefinitely.", "kind": "method", "line": 285, "name": "make_loop_detector", "signature": "def make_loop_detector(window, min_repeats)"}, {"doc": "Return True if the per-token wall time exceeds `per_token_s`.\nUseful for catching token-generation stalls (rare on CPU, but\nhappens under memory pressure).", "kind": "method", "line": 314, "name": "make_timeout_hook", "signature": "def make_timeout_hook(per_token_s)"}, {"kind": "method", "line": 56, "name": "__init__", "signature": "def __init__(self)"}, {"kind": "method", "line": 62, "name": "put", "signature": "def put(self, tok)"}, {"kind": "method", "line": 67, "name": "mark_done", "signature": "def mark_done(self)"}, {"doc": "Return all tokens emitted so far, atomic snapshot.", "kind": "method", "line": 72, "name": "drain", "signature": "def drain(self)"}, {"doc": "Block up to `timeout` seconds for a new token. Returns True\nif a new token arrived (or stream closed), False on timeout.", "kind": "method", "line": 77, "name": "wait_for_new", "signature": "def wait_for_new(self, timeout)"}, {"kind": "method", "line": 86, "name": "is_closed", "signature": "def is_closed(self)"}, {"kind": "method", "line": 90, "name": "__len__", "signature": "def __len__(self)"}, {"kind": "method", "line": 117, "name": "__post_init__", "signature": "def __post_init__(self)"}, {"kind": "method", "line": 156, "name": "__init__", "signature": "def __init__(self, model, ctx, stream, max_new_tokens, temperature, top_k, repetition_penalty, max_seq_len)"}, {"doc": "Asynchronously stop the generation. Safe to call from any\nthread (e.g. a watchdog thread or the main UI loop).", "kind": "method", "line": 177, "name": "cancel", "signature": "def cancel(self)"}, {"kind": "method", "line": 182, "name": "_should_cancel", "signature": "def _should_cancel(self)"}, {"doc": "Execute the generation loop. Returns when the model emits\nEOS, hits max_new_tokens, a hook returns True, or cancel() is\ncalled.", "kind": "method", "line": 185, "name": "run", "signature": "def run(self, stop_hooks)"}, {"kind": "method", "line": 292, "name": "hook", "signature": "def hook(generated)"}, {"kind": "method", "line": 320, "name": "hook", "signature": "def hook(generated)"}]}, {"doc": "Smoke test for eval.governor (TokenStream + GenerationGovernor).  Verifies: - TokenStream threadsafety with a producer/consumer scenario - GenerationGovernor emits one StopReason per call - Loop detector actually fires - User cancel() works", "id": "eval/governor_smoke.py", "kind": "module", "label": "governor_smoke.py", "language": "py", "sha256": "6adca837aa8182c8", "symbol_count": 7, "symbols": [{"kind": "function", "line": 30, "name": "load_model", "signature": "def load_model()"}, {"kind": "function", "line": 49, "name": "test_tokenstream_threadsafety", "signature": "def test_tokenstream_threadsafety()"}, {"kind": "function", "line": 79, "name": "test_governor_basic", "signature": "def test_governor_basic()"}, {"kind": "function", "line": 98, "name": "test_loop_detector", "signature": "def test_loop_detector()"}, {"kind": "function", "line": 118, "name": "test_cancel", "signature": "def test_cancel()"}, {"kind": "function", "line": 53, "name": "producer", "signature": "def producer()"}, {"kind": "function", "line": 59, "name": "consumer", "signature": "def consumer()"}]}, {"doc": "Harness for evaluating TopoGPT3 on HumanEval (164 problems).  Faithful to the official HumanEval protocol: for each problem we feed the model the function signature and docstring, let it produce a completion, extract the candidate function (everything from `def` up to a sentinel), and run the hidden test against it. We do NOT use `entry_point` from the dataset because the prompt we feed the model already contains it.  Two sampler modes are supported: - \"standard\" -> topogpt3.InferencePipeline - \"hrm\"      -> topogpt3.HRMInferencePipeline  Results are written to JSONL so multiple sampler configurations can share a single HumanEval cache and be compared later.", "id": "eval/harness.py", "kind": "module", "label": "harness.py", "language": "py", "sha256": "384baed586db5396", "symbol_count": 12, "symbols": [{"kind": "function", "line": 59, "name": "load_humaneval", "signature": "def load_humaneval(cache_dir)"}, {"doc": "Return the exact prompt text fed to the model.\n\nHumanEval's `prompt` field already contains the function signature and\ndocstring, with the body to be completed starting on the next line.", "kind": "function", "line": 75, "name": "build_prompt", "signature": "def build_prompt(problem)"}, {"doc": "Combine prompt + completion into a single Python source string.\n\nThe completion may itself start with whitespace/indentation that\nbelongs inside the function body. We strip leading blank lines and\nthen concatenate; we also stop at the first top-level `def ` or\n`class ` to avoid the model continuing with extra functions.\n\nRobustness fixes:\n  - Strip the special <|endoftext|> (GPT-2 EOT) token that the model\n    emits at the end of every generation. Leaving it in the candidate\n    produces a SyntaxError and zeroes the pass rate.\n  - Drop any training-format delimiters (### Response, <|assistant|>,\n    <|user|>) that leak from the instruction-tuning corpus.\n  - Cut at the first top-level def/class/__main__ guard after the\n    function body has started.", "kind": "function", "line": 100, "name": "extract_candidate", "signature": "def extract_candidate(prompt, completion)"}, {"doc": "Execute the candidate against the hidden test.\n\nReturns (passed, message, stdout, stderr, traceback). We follow HumanEval's\n`evaluate` function: build namespace, exec the candidate, exec the test,\nexpect `check(candidate) == None`.", "kind": "function", "line": 150, "name": "run_one_test", "signature": "def run_one_test(problem, candidate_src, timeout)"}, {"doc": "Sandboxed variant of `run_one_test`. Runs the candidate in a\nsubprocess with stripped builtins, AST pre-check, and OS-enforced\ntimeout. Drop-in replacement: same 5-tuple return.\n\nEnable by passing `--sandbox` to `harness.py` (not yet wired) or\nby calling this function directly from your own evaluation script.", "kind": "function", "line": 172, "name": "run_one_test_sandboxed", "signature": "def run_one_test_sandboxed(problem, candidate_src, timeout, sandbox_cfg)"}, {"doc": "Backwards-compatible shim. The real implementation lives in\n`eval.samplers` as a decorator-based registry. We re-export here\nso existing imports of `from eval.harness import make_sampler`\nkeep working. New code should import from `eval.samplers`.", "kind": "function", "line": 195, "name": "make_sampler", "signature": "def make_sampler(mode, settings_kwargs)"}, {"doc": "Run a single completion and return (raw_output_text, metrics_dict).", "kind": "function", "line": 204, "name": "completion_for_problem", "signature": "def completion_for_problem(sampler, prompt)"}, {"doc": "Build the model and tokenizer once, run many generations.", "kind": "class", "line": 217, "name": "ModelLoader", "signature": "class ModelLoader"}, {"kind": "method", "line": 272, "name": "evaluate_problem", "signature": "def evaluate_problem(problem, loader, args, sample_idx)"}, {"kind": "method", "line": 315, "name": "main", "signature": "def main()"}, {"kind": "method", "line": 220, "name": "__init__", "signature": "def __init__(self, ckpt_dir, ckpt_name, device)"}, {"kind": "method", "line": 246, "name": "generate", "signature": "def generate(self, prompt, max_new_tokens, temperature, top_k, repetition_penalty)"}]}, {"doc": "Hodge-CM offline ablation on real TopoGPT3 checkpoint. Runs on GPU when available (--device cuda), else CPU. No datasets removed: uses the real checkpoint kernels (all 12) + real layer weights; nothing subsampled, no tier dropped. Saves JSON to eval/runs/hodge_cm_ablation.json Covers: Fase1 diag + offline ablation (CM/Hodge/RAND/HECKE) + functional probe.", "id": "eval/hodge_cm_ablation.py", "kind": "module", "label": "hodge_cm_ablation.py", "language": "py", "sha256": "b125fb7e2568003c", "symbol_count": 6, "symbols": [{"kind": "function", "line": 28, "name": "resolve_device", "signature": "def resolve_device(req)"}, {"kind": "function", "line": 48, "name": "mat_stats", "signature": "def mat_stats(kr, ki, r)"}, {"kind": "function", "line": 64, "name": "heat_smooth", "signature": "def heat_smooth(kr, ki, t, device)"}, {"kind": "function", "line": 74, "name": "main", "signature": "def main()"}, {"kind": "function", "line": 70, "name": "sm", "signature": "def sm(x)"}, {"kind": "function", "line": 141, "name": "mean", "signature": "def mean(k, m)"}]}, {"doc": "End-to-end smoke test of all P0+P1 components working together.  Verifies that `run_one_test_sandboxed` can run a *valid* HumanEval candidate through the sandbox and get a pass=True result. This exercises the integration of: - sandbox.py (P0) - harness.py integration (the new run_one_test_sandboxed) - HumanEval canonical protocol (prompt + completion + test)", "id": "eval/integration_smoke.py", "kind": "module", "label": "integration_smoke.py", "language": "py", "sha256": "d23b6c8473c4cf4b", "symbol_count": 1, "symbols": [{"kind": "function", "line": 18, "name": "main", "signature": "def main()"}]}, {"doc": "Analisis post-hoc del noise sweep.  Generaciones del MISMO prompt bajo distintos niveles de ruido -> comparar con metricas que NO son pass@1 (porque los problemas triviales saturan):  - generation_exact_match:  % de generaciones que matchean exactamente el baseline (token por token) - prefix_agreement@50:    % de pares (baseline, noisy) que comparten el mismo prefijo de 50 tokens - levenshtein_dist:       distancia de edicion normalizada al baseline - token_jaccard:          interseccion / union de tokens generados - bleu_1:                 unigrama precision - syntax_ok_rate:         % que pasa ast.parse (sintaxis Python valida)  Salida: eval/runs/noise_analysis_<tag>.json", "id": "eval/noise_analysis.py", "kind": "module", "label": "noise_analysis.py", "language": "py", "sha256": "ab6bbc133fa88eb1", "symbol_count": 3, "symbols": [{"kind": "function", "line": 43, "name": "_load", "signature": "def _load(p)"}, {"doc": "Para cada problema, mira si pasa consistentemente a traves de los\n4 niveles de ruido. Devuelve:\n  - always_pass, always_fail, mixed (count)\n  - per_sigma_pass_lists: {sigma: {task_id: bool}}", "kind": "function", "line": 47, "name": "consistency_across_runs", "signature": "def consistency_across_runs(per_run)"}, {"kind": "function", "line": 83, "name": "main", "signature": "def main()"}]}, {"doc": "Barrido de ruido en los pesos espectrales del checkpoint TopoGPT3.  Para cada nivel sigma en --sigmas: 1. Carga el checkpoint base (NO modifica el archivo, solo los pesos en RAM) 2. Inyecta ruido gaussiano N(0, sigma) en los kernels espectrales cuaternionicos (kr_w/x/y/z, ki_w/x/y/z) y solo en ellos. Asi aislamos el efecto del ruido sobre la parte que el marco teorico dice que esta protegida topologicamente. 3. Genera pass@1 (greedy, T=0) sobre los primeros N problemas de HumanEval (subset para mantener tiempo de pared manejable) 4. Ejecuta los tests canonicos y mide pass rate  Salida: eval/runs/noise_<sigma>_<tag>.jsonl eval/runs/noise_sweep_<timestamp>.jsonl (resumen agregado)", "id": "eval/noise_sweep.py", "kind": "module", "label": "noise_sweep.py", "language": "py", "sha256": "0cbc44aa54b2844b", "symbol_count": 4, "symbols": [{"doc": "Anade N(0, sigma) a TODOS los kernels espectrales (kr_*, ki_*).\nRetorna un dict con conteo de tensores ruidosos y de parametros\nmodificados.", "kind": "function", "line": 46, "name": "inject_noise", "signature": "def inject_noise(model, sigma, seed)"}, {"doc": "Reconstruye TopoGPT2 alineado con el checkpoint, sin acceso a\nharness.ModelLoader (queremos un loader limpio que no comparta\nestado con corridas paralelas).", "kind": "function", "line": 74, "name": "load_model", "signature": "def load_model(ckpt_dir, ckpt_name, device)"}, {"kind": "function", "line": 99, "name": "generate_one", "signature": "def generate_one(model, tok, prompt, max_new_tokens, device)"}, {"kind": "function", "line": 117, "name": "main", "signature": "def main()"}]}, {"doc": "Self-repair loop on top of a greedy JSONL.  Takes the failed problems from --input, builds a rejection-feedback prompt that contains: - the original HumanEval prompt (signature + docstring) - the candidate the model wrote on its first attempt - the traceback from the hidden test - a \"# fix:\" cue  and re-prompts the model to rewrite the function. Runs N rounds. Each problem's *best* outcome across rounds is recorded.  Output: a new JSONL with the same shape as harness.py.", "id": "eval/repair.py", "kind": "module", "label": "repair.py", "language": "py", "sha256": "826f80ea70d53790", "symbol_count": 6, "symbols": [{"kind": "function", "line": 36, "name": "_new_loader", "signature": "def _new_loader(ckpt_dir, ckpt_name)"}, {"kind": "function", "line": 49, "name": "extract_candidate", "signature": "def extract_candidate(prompt, completion)"}, {"kind": "function", "line": 75, "name": "run_test", "signature": "def run_test(problem, candidate_src)"}, {"kind": "function", "line": 89, "name": "build_repair_prompt", "signature": "def build_repair_prompt(prompt, candidate, err, entry_point)"}, {"kind": "function", "line": 104, "name": "gen", "signature": "def gen(model, tok, text, max_new_tokens, temperature, top_k, rep_penalty)"}, {"kind": "function", "line": 119, "name": "main", "signature": "def main()"}]}, {"doc": "Aggregate every JSONL in eval/runs into a final report.  Reads runs from the original pass@k runs, the HRM run, the repair run, and produces: - pass@1 / pass@k tables - error-class breakdowns - wall-clock / throughput - a comparison standard vs HRM - a self-repair impact summary - emits REPORT.md next to the runs/", "id": "eval/report.py", "kind": "module", "label": "report.py", "language": "py", "sha256": "87956c8496bab9b5", "symbol_count": 6, "symbols": [{"kind": "function", "line": 25, "name": "pass_at_k", "signature": "def pass_at_k(n, c, k)"}, {"kind": "function", "line": 31, "name": "classify_error", "signature": "def classify_error(msg)"}, {"kind": "function", "line": 52, "name": "load_jsonl", "signature": "def load_jsonl(p)"}, {"kind": "function", "line": 56, "name": "summarize_run", "signature": "def summarize_run(p)"}, {"kind": "function", "line": 90, "name": "repair_summary", "signature": "def repair_summary(repair_path, baseline_path)"}, {"kind": "function", "line": 117, "name": "main", "signature": "def main()"}]}, {"doc": "Registry of sampler constructors for the HumanEval harness.  Replaces the hardcoded `if mode == \"standard\": ... elif mode == \"hrm\": ...` chain in `eval.harness.make_sampler` with a decorator-based registry that mirrors the pattern in `claude-code-main/src/tools.ts`.  The pattern: - `@register_sampler(\"name\")` decorates a factory function that takes a `settings_kwargs` dict (already filtered for sampler-specific keys) and returns an object with a `run()` method (or just a sampler object that `harness.evaluate_problem` knows how to drive). - `build_sampler(\"name\", settings_kwargs)` is the public entry point. - `list_samplers()` returns the registered names for `--help` output.  Adding a new sampler is then a one-decorator change, not an edit to the harness's control flow.", "id": "eval/samplers.py", "kind": "module", "label": "samplers.py", "language": "py", "sha256": "60f441ded1682d30", "symbol_count": 7, "symbols": [{"doc": "Decorator. Register a factory under `name`. If `enabled_env` is set,\nthe factory is only registered when that env var is truthy. This\nmirrors the `feature('XXX')` gating in claude-code-main/src/tools.ts.", "kind": "function", "line": 36, "name": "register_sampler", "signature": "def register_sampler(name)"}, {"kind": "function", "line": 55, "name": "_is_env_truthy", "signature": "def _is_env_truthy(name)"}, {"kind": "function", "line": 64, "name": "_make_standard", "signature": "def _make_standard(settings_kwargs)"}, {"kind": "function", "line": 69, "name": "_make_hrm", "signature": "def _make_hrm(settings_kwargs)"}, {"kind": "function", "line": 86, "name": "list_samplers", "signature": "def list_samplers()"}, {"doc": "Construct a sampler. Drop-in replacement for the old\n`make_sampler(mode, settings_kwargs)` in `eval.harness`.", "kind": "function", "line": 90, "name": "build_sampler", "signature": "def build_sampler(mode, settings_kwargs)"}, {"kind": "function", "line": 42, "name": "deco", "signature": "def deco(fn)"}]}, {"doc": "Sandbox for executing model-generated code during HumanEval evaluation.  This module replaces the bare `exec()` call in `eval/harness.py:run_one_test` with a defence-in-depth check inspired by Claude Code's BashTool permission gates (see `claude-code-main/src/tools/BashTool/bashSecurity.ts`).  The threat model: - A language model emits Python source as a \"candidate function\". - The candidate is `exec()`'d alongside a hidden test. - Without protection, the model can `import os; os.system('rm -rf /')`, read secrets, fork-bomb, or hang the evaluator forever.  Layered defences (each can be disabled independently for debugging): 1. AST pre-check: parse the candidate, reject anything that imports dangerous modules, calls dangerous builtins, or shadows `__builtins__`. 2. Builtin whitelist: even if the candidate parses, `safe_exec` provides a stripped `__builtins__` without `open`, `exec`, `eval`, `__import__`, `compile`, `getattr` (controversial but standard). 3. Subprocess isolation: `safe_exec` runs the program in a child process so the OS enforces the timeout (vs. signal-based which the main thread can swallow). 4. Output capture: stdout/stderr are piped, not inherited from the parent terminal.  Usage: from eval.sandbox import safe_exec, check_safety, SandboxConfig  cfg = SandboxConfig(timeout=10.0, dry_run=False) ok, reason = check_safety(candidate_src, cfg) if not ok:", "id": "eval/sandbox.py", "kind": "module", "label": "sandbox.py", "language": "py", "sha256": "bf933df740f00cef", "symbol_count": 9, "symbols": [{"doc": "One knob per defence layer. Defaults match HumanEval-style eval.", "kind": "class", "line": 53, "name": "SandboxConfig", "signature": "class SandboxConfig"}, {"doc": "Return the set of top-level names brought into scope by imports.", "kind": "method", "line": 100, "name": "_names_imported", "signature": "def _names_imported(tree)"}, {"doc": "Find Attribute nodes whose attr is in `blocked`. Returns attr names found.", "kind": "method", "line": 114, "name": "_blocked_dunder_access", "signature": "def _blocked_dunder_access(tree, blocked)"}, {"doc": "Compute max nesting depth of the AST. Catches obfuscated huge trees.", "kind": "method", "line": 123, "name": "_max_depth", "signature": "def _max_depth(tree)"}, {"doc": "Return (ok, reason). `reason` is \"\" when ok, else a human-readable\none-line explanation. Reasons are stable (used in test fixtures).", "kind": "method", "line": 133, "name": "check_safety", "signature": "def check_safety(source, cfg)"}, {"kind": "method", "line": 254, "name": "_build_worker_src", "signature": "def _build_worker_src(allowed_builtin_names, program_src, blocked_modules)"}, {"doc": "Execute `program_src` in a sandboxed child process. Returns the same\n5-tuple as `eval.harness.run_one_test` for drop-in compatibility.\n\nThe child is killed (SIGKILL) by the OS after `cfg.timeout` seconds.", "kind": "method", "line": 270, "name": "safe_exec", "signature": "def safe_exec(program_src, cfg, extra_globals)"}, {"kind": "method", "line": 373, "name": "describe_policy", "signature": "def describe_policy(cfg)"}, {"kind": "method", "line": 125, "name": "d", "signature": "def d(node, cur)"}]}, {"doc": "Smoke test for eval.sandbox.  Verifies all four defence layers: L1 (AST pre-check): blocked imports & dunder attrs are rejected L2 (builtin whitelist): open/exec/etc raise NameError in the child L3 (subprocess isolation): infinite loops are killed at OS level L4 (output capture): stdout/stderr from the candidate are returned", "id": "eval/sandbox_smoke.py", "kind": "module", "label": "sandbox_smoke.py", "language": "py", "sha256": "89cee69fb565fbe9", "symbol_count": 1, "symbols": [{"kind": "function", "line": 15, "name": "main", "signature": "def main()"}]}, {"doc": "Smoke test: load the TopoGPT3 checkpoint and produce a small completion.  Used as the first gate: if this fails we abort HumanEval.", "id": "eval/smoke.py", "kind": "module", "label": "smoke.py", "language": "py", "sha256": "94ee21110f91e4c2", "symbol_count": 2, "symbols": [{"kind": "function", "line": 17, "name": "run_standard", "signature": "def run_standard()"}, {"kind": "function", "line": 36, "name": "run_hrm", "signature": "def run_hrm()"}]}, {"doc": "Barrido de temperatura x top-k sobre HumanEval.  Mide pass@1 en modo greedy (T=0) y pass@5 a temperaturas crecientes para mapear la \"fase\" de generacion del modelo:  - cristal: pass@1 alto, poca varianza entre samples - vidrio:  pass@1 bajo, alta varianza - caotico: pass@1 ~= 0, alta diversidad pero sin aciertos  Salida: eval/runs/temp_sweep_<tag>.jsonl (resumen) eval/runs/temp_<T>_top<k>_<tag>.jsonl (detalle por config)", "id": "eval/temp_sweep.py", "kind": "module", "label": "temp_sweep.py", "language": "py", "sha256": "61dea99f11f6e36b", "symbol_count": 5, "symbols": [{"kind": "function", "line": 39, "name": "generate_one", "signature": "def generate_one(model, tok, prompt, max_new_tokens, temperature, top_k, device, seed_offset)"}, {"kind": "function", "line": 58, "name": "evaluate_problems", "signature": "def evaluate_problems(model, tok, problems, max_new_tokens, temperature, top_k, n_samples, device)"}, {"kind": "function", "line": 88, "name": "pass_at_k_unbiased", "signature": "def pass_at_k_unbiased(n, c, k)"}, {"kind": "function", "line": 96, "name": "summarize", "signature": "def summarize(results, n_samples)"}, {"kind": "function", "line": 116, "name": "main", "signature": "def main()"}]}, {"doc": "Standalone inference script for TopoExploit.  Loads the 125M TopoGPT-2 model and performs inference on exploit-related tasks using prompts derived from the .exploitgym_repo dataset.  Modes: --prompt \"text\"       Single prompt inference --eval-holdout        Evaluate on holdout tasks from the repo --interactive         Interactive prompt REPL  Tiers: vulnerability_analysis, patch_analysis, exploit_development", "id": "infer_exploitgym.py", "kind": "module", "label": "infer_exploitgym.py", "language": "py", "sha256": "59c2cf4e815e3c17", "symbol_count": 11, "symbols": [{"kind": "function", "line": 41, "name": "load_task_ids", "signature": "def load_task_ids()"}, {"doc": "Return a dict with keys: task_id, family, name, description, patch, pov.", "kind": "function", "line": 49, "name": "load_task", "signature": "def load_task(task_id)"}, {"kind": "function", "line": 122, "name": "build_prompt", "signature": "def build_prompt(task_info, tier)"}, {"doc": "Load tokenizer, build TopoGPT-2, apply Gauss patch, load weights.", "kind": "function", "line": 136, "name": "load_model", "signature": "def load_model(ckpt_dir, ckpt_name, device)"}, {"doc": "Autoregressive generation with streaming + n-gram repetition blocking.\n\nGenerates token by token using the model's KV cache. When ``stream`` is\nTrue, decoded text is flushed incrementally (via ``stream_cb`` if given,\notherwise to stdout) so the user sees output as it is produced instead of\nwaiting for the whole sequence. Repeated n-grams of size\n``no_repeat_ngram`` are suppressed to avoid degenerate repetition loops.", "kind": "function", "line": 177, "name": "generate", "signature": "def generate(model, tokenizer, prompt)"}, {"kind": "function", "line": 257, "name": "run_single_prompt", "signature": "def run_single_prompt(args, model, tokenizer)"}, {"kind": "function", "line": 274, "name": "run_eval_holdout", "signature": "def run_eval_holdout(args, model, tokenizer)"}, {"kind": "function", "line": 331, "name": "run_interactive", "signature": "def run_interactive(args, model, tokenizer)"}, {"kind": "function", "line": 380, "name": "parse_args", "signature": "def parse_args()"}, {"kind": "function", "line": 420, "name": "main", "signature": "def main()"}, {"kind": "function", "line": 205, "name": "sample", "signature": "def sample(logits, seen_ids)"}]}, {"id": "install.sh", "kind": "module", "label": "install.sh", "language": "sh", "sha256": "c907d80fd6734993", "symbol_count": 0, "symbols": []}, {"doc": "TopoExploit: 125M params trained on ExploitGym  Usage: ./run_exploitgym.sh prepare     -- clone repo + tokenize data ./run_exploitgym.sh transfer    -- transfer weights from small model ./run_exploitgym.sh train       -- train from scratch ./run_exploitgym.sh train-init  -- train with transferred weights ./run_exploitgym.sh eval        -- evaluate on holdout ./run_exploitgym.sh full        -- prepare + transfer + train + eval", "id": "run_exploitgym.sh", "kind": "module", "label": "run_exploitgym.sh", "language": "sh", "sha256": "ee4c18d3dfcd5033", "symbol_count": 0, "symbols": []}, {"id": "run_exploitgym_v2.sh", "kind": "module", "label": "run_exploitgym_v2.sh", "language": "sh", "sha256": "2a87d3e092ef176b", "symbol_count": 4, "symbols": [{"kind": "function", "line": 24, "name": "banner"}, {"kind": "function", "line": 32, "name": "train_v2"}, {"kind": "function", "line": 45, "name": "eval_v2"}, {"kind": "function", "line": 54, "name": "infer_v2"}]}, {"doc": "Hodge-CM offline ablation runner (GPU si el build torch lo permite, si no CPU). No recorta nada: 12 kernels reales + probe funcional + control vanilla con curriculum completo (codealpaca/code_feedback/magicoder/tiny_the_stack, train/val/holdout disjuntos). Sin sintetico silencioso. Usage: ./run_hodge_cm_ablation.sh [--device auto|cuda|cpu]", "id": "run_hodge_cm_ablation.sh", "kind": "module", "label": "run_hodge_cm_ablation.sh", "language": "sh", "sha256": "dc5d35e6927bdcd7", "symbol_count": 0, "symbols": []}, {"id": "run_merged.sh", "kind": "module", "label": "run_merged.sh", "language": "sh", "sha256": "3a9bb5f3ab409d03", "symbol_count": 7, "symbols": [{"kind": "function", "line": 33, "name": "kill_gpu_processes"}, {"kind": "function", "line": 63, "name": "restore_gpu_processes"}, {"kind": "function", "line": 84, "name": "banner"}, {"kind": "function", "line": 93, "name": "prepare_data"}, {"kind": "function", "line": 107, "name": "train_merged"}, {"kind": "function", "line": 119, "name": "eval_merged"}, {"kind": "function", "line": 128, "name": "infer_merged"}]}, {"doc": "Synthetic Dataset Generator for TopoGPT2.  Generates high-quality code instruction-tuning data from existing source files using a multi-stage LLM pipeline:  file → analysis → spec → chain-of-thought → vague question → JSONL  Each sample (JSONL line) contains: { \"instruction\":  \"vague natural question\", \"thinking\":     \"chain-of-thought reasoning\", \"spec\":         \"detailed spec-driven prompt\", \"todo\":         [\"task 1\", \"task 2\", ...], \"response\":     \"```language\\noriginal clean code\\n```\", \"file_path\":    \"src/foo/bar.py\", \"lang\":         \"python\", \"checksum\":     \"sha256 of original code\", }  Pipeline is designed for efficiency: - One LLM call per file (master prompt, one-shot) - Streaming JSONL writes (never holds full dataset in memory) - SHA256 dedup across the full corpus - Resumable: tracks processed files in a manifest - Batch-friendly: process N files per run  Backend: Groq API (Llama-3.3-70B, fastest/cheapest) or OpenRouter.", "id": "synthetic_dataset.py", "kind": "module", "label": "synthetic_dataset.py", "language": "py", "sha256": "a91dbe7020222e8f", "symbol_count": 35, "symbols": [{"doc": "Abstract LLM backend. Subclass for each provider.", "kind": "class", "line": 61, "name": "LLMBackend", "signature": "class LLMBackend"}, {"doc": "Groq API backend using requests.\n\nSupports models: llama-3.3-70b-versatile, deepseek-r1.\nSet GROQ_API_KEY env var.", "kind": "class", "line": 71, "name": "GroqBackend", "signature": "class GroqBackend(LLMBackend)"}, {"doc": "OpenRouter unified API backend.\n\nSupports any OpenRouter model:\n    anthropic/claude-3.5-sonnet,\n    openai/gpt-4o,\n    deepseek/deepseek-chat,\n    google/gemini-2.0-flash-thinking,\nSet OPENROUTER_API_KEY env var.", "kind": "class", "line": 121, "name": "OpenRouterBackend", "signature": "class OpenRouterBackend(LLMBackend)"}, {"doc": "Ollama local inference backend.\n\nSupports any local model: llama3.1:8b, granite4.1:3b, etc.\nConnects to Ollama server at OLLAMA_HOST (default: http://localhost:11434).", "kind": "class", "line": 177, "name": "OllamaBackend", "signature": "class OllamaBackend(LLMBackend)"}, {"doc": "Factory for LLM backends.", "kind": "method", "line": 227, "name": "build_backend", "signature": "def build_backend(provider, model)"}, {"doc": "Validate that a generated sample meets quality bar.\n\nReturns (is_valid, reason).", "kind": "method", "line": 330, "name": "validate_sample", "signature": "def validate_sample(sample)"}, {"doc": "Tracks processed files for resumability.", "kind": "class", "line": 364, "name": "ProcessedManifest", "signature": "class ProcessedManifest"}, {"doc": "Generates synthetic instruction-tuning data from source files.\n\nPipeline (one LLM call per file):\n    file → MASTER_PROMPT → LLM → validate → dedup → JSONL\n\nFeatures:\n- Streaming JSONL writes (bounded RAM)\n- SHA256 dedup across corpus\n- Resumable (manifest tracks progress)\n- Threaded request batching for throughput\n- Configurable quality thresholds", "kind": "class", "line": 399, "name": "SyntheticDatasetGenerator", "signature": "class SyntheticDatasetGenerator"}, {"kind": "method", "line": 614, "name": "build_logger", "signature": "def build_logger(level)"}, {"kind": "method", "line": 625, "name": "parse_args", "signature": "def parse_args()"}, {"doc": "Load file paths from CLI args or file.", "kind": "method", "line": 652, "name": "load_paths", "signature": "def load_paths(paths_arg, paths_file, max_files)"}, {"kind": "method", "line": 667, "name": "main", "signature": "def main()"}, {"kind": "method", "line": 64, "name": "generate", "signature": "def generate(self, prompt)"}, {"kind": "method", "line": 67, "name": "name", "signature": "def name(self)"}, {"kind": "method", "line": 78, "name": "__init__", "signature": "def __init__(self, model, api_key, max_tokens, temperature, timeout)"}, {"kind": "method", "line": 95, "name": "name", "signature": "def name(self)"}, {"kind": "method", "line": 98, "name": "generate", "signature": "def generate(self, prompt)"}, {"kind": "method", "line": 132, "name": "__init__", "signature": "def __init__(self, model, api_key, max_tokens, temperature, timeout)"}, {"kind": "method", "line": 151, "name": "name", "signature": "def name(self)"}, {"kind": "method", "line": 154, "name": "generate", "signature": "def generate(self, prompt)"}, {"kind": "method", "line": 184, "name": "__init__", "signature": "def __init__(self, model, host, max_tokens, temperature, timeout)"}, {"kind": "method", "line": 198, "name": "name", "signature": "def name(self)"}, {"kind": "method", "line": 201, "name": "generate", "signature": "def generate(self, prompt)"}, {"kind": "method", "line": 374, "name": "load", "signature": "def load(path)"}, {"kind": "method", "line": 387, "name": "save", "signature": "def save(self, path)"}, {"kind": "method", "line": 418, "name": "__init__", "signature": "def __init__(self, backend, output_path, manifest_path, logger, max_workers, max_file_chars)"}, {"doc": "Background thread that drains the queue and writes JSONL lines.", "kind": "method", "line": 447, "name": "_jsonl_writer", "signature": "def _jsonl_writer(self)"}, {"kind": "method", "line": 465, "name": "_enqueue_sample", "signature": "def _enqueue_sample(self, sample)"}, {"kind": "method", "line": 468, "name": "_flush_writer", "signature": "def _flush_writer(self)"}, {"doc": "Read file content and detect language. Truncate if needed.", "kind": "method", "line": 477, "name": "_read_file", "signature": "def _read_file(self, path)"}, {"kind": "method", "line": 490, "name": "_build_prompt", "signature": "def _build_prompt(self, content, lang)"}, {"doc": "Call LLM with retry logic.", "kind": "method", "line": 496, "name": "_generate_sample", "signature": "def _generate_sample(self, content, lang)"}, {"doc": "Process a single file. Returns True if a sample was written.", "kind": "method", "line": 533, "name": "process_file", "signature": "def process_file(self, path)"}, {"doc": "Process a batch of files in parallel using thread pool.", "kind": "method", "line": 568, "name": "process_batch", "signature": "def process_batch(self, paths)"}, {"doc": "Signal end of processing and flush writer.", "kind": "method", "line": 590, "name": "finish", "signature": "def finish(self)"}]}, {"id": "tests/test_jlens.py", "kind": "module", "label": "test_jlens.py", "language": "py", "sha256": "3fd9fb4a583778eb", "symbol_count": 51, "symbols": [{"doc": "Feature: valid_position_mask excludes attention-sink and final positions.", "kind": "class", "line": 17, "name": "TestValidPositionMask", "signature": "class TestValidPositionMask"}, {"doc": "Feature: jacobian_for_prompt computes J_l for one prompt.", "kind": "class", "line": 52, "name": "TestJacobianForPrompt", "signature": "class TestJacobianForPrompt"}, {"doc": "Feature: fit() averages Jacobians over multiple prompts.", "kind": "class", "line": 172, "name": "TestFit", "signature": "class TestFit"}, {"doc": "Feature: JacobianLens saves, loads, applies, and merges.", "kind": "class", "line": 210, "name": "TestJacobianLens", "signature": "class TestJacobianLens"}, {"doc": "Feature: fit() with checkpoint resume works correctly.", "kind": "class", "line": 367, "name": "TestFitCheckpoint", "signature": "class TestFitCheckpoint"}, {"doc": "Feature: Config classes centralize all tunable parameters.", "kind": "class", "line": 473, "name": "TestConfig", "signature": "class TestConfig"}, {"doc": "Feature: Application config controls readout behavior.", "kind": "class", "line": 494, "name": "TestTopoGPT3JLensAppConfig", "signature": "class TestTopoGPT3JLensAppConfig"}, {"doc": "Scenario: Correct mask for a standard-length prompt.", "kind": "method", "line": 20, "name": "test_basic_mask", "signature": "def test_basic_mask(self)"}, {"doc": "Scenario: Too-short prompt raises ValueError.", "kind": "method", "line": 29, "name": "test_too_short_raises", "signature": "def test_too_short_raises(self)"}, {"doc": "Scenario: Negative skip_first raises ValueError.", "kind": "method", "line": 34, "name": "test_negative_skip_raises", "signature": "def test_negative_skip_raises(self)"}, {"doc": "Scenario: skip_first=0 includes all but final position.", "kind": "method", "line": 39, "name": "test_all_positions_valid", "signature": "def test_all_positions_valid(self)"}, {"doc": "Scenario: Exact minimum length (skip_first + 2) works.", "kind": "method", "line": 45, "name": "test_exact_minimum_length", "signature": "def test_exact_minimum_length(self)"}, {"kind": "method", "line": 56, "name": "model", "signature": "def model(self)"}, {"doc": "Scenario: Returns Jacobians for all requested source layers.", "kind": "method", "line": 63, "name": "test_returns_jacobians_for_source_layers", "signature": "def test_returns_jacobians_for_source_layers(self, model)"}, {"doc": "Scenario: J_{n_layers-2} has diag ~= 1 (identity property).", "kind": "method", "line": 76, "name": "test_late_layer_jacobian_close_to_identity", "signature": "def test_late_layer_jacobian_close_to_identity(self, model)"}, {"doc": "Scenario: Earlier layers compound deviations from identity.", "kind": "method", "line": 85, "name": "test_earlier_layers_further_from_identity", "signature": "def test_earlier_layers_further_from_identity(self, model)"}, {"doc": "Scenario: J_{n_layers-2} equals I + W_{last} exactly.\n\nFor TinyDecoder with block = h + 0.1*W*h, J_{n_layers-2} = I + W.", "kind": "method", "line": 95, "name": "test_exact_jacobian_for_last_block", "signature": "def test_exact_jacobian_for_last_block(self, model)"}, {"doc": "Scenario: Negative layer indices are normalized correctly.", "kind": "method", "line": 110, "name": "test_negative_layer_indices", "signature": "def test_negative_layer_indices(self, model)"}, {"doc": "Scenario: Out-of-range layers raise ValueError.", "kind": "method", "line": 133, "name": "test_out_of_range_layers_rejected", "signature": "def test_out_of_range_layers_rejected(self, model)"}, {"doc": "Scenario: source_layers must be below target_layer.", "kind": "method", "line": 145, "name": "test_source_below_target_enforced", "signature": "def test_source_below_target_enforced(self, model)"}, {"doc": "Scenario: target_layer out of range raises ValueError.", "kind": "method", "line": 158, "name": "test_target_out_of_range_raises", "signature": "def test_target_out_of_range_raises(self, model)"}, {"kind": "method", "line": 176, "name": "model", "signature": "def model(self)"}, {"doc": "Scenario: fit() returns JacobianLens with correct metadata.", "kind": "method", "line": 183, "name": "test_fit_returns_lens_with_correct_attributes", "signature": "def test_fit_returns_lens_with_correct_attributes(self, model)"}, {"doc": "Scenario: No valid prompts raises ValueError.", "kind": "method", "line": 191, "name": "test_fit_empty_prompts_raises", "signature": "def test_fit_empty_prompts_raises(self, model)"}, {"doc": "Scenario: Too-short prompts are skipped.", "kind": "method", "line": 196, "name": "test_fit_skips_short_prompts", "signature": "def test_fit_skips_short_prompts(self, model)"}, {"doc": "Scenario: Default source_layers covers all layers below target.", "kind": "method", "line": 202, "name": "test_fit_with_default_source_layers", "signature": "def test_fit_with_default_source_layers(self, model)"}, {"kind": "method", "line": 214, "name": "model", "signature": "def model(self)"}, {"kind": "method", "line": 222, "name": "fitted_lens", "signature": "def fitted_lens(self, model)"}, {"doc": "Scenario: save/load preserves jacobians (fp16 tolerance).", "kind": "method", "line": 226, "name": "test_save_and_load_round_trip", "signature": "def test_save_and_load_round_trip(self, fitted_lens, tmp_path)"}, {"doc": "Scenario: apply() returns correct logit shapes.", "kind": "method", "line": 242, "name": "test_apply_returns_correct_shapes", "signature": "def test_apply_returns_correct_shapes(self, fitted_lens, model)"}, {"doc": "Scenario: Transported late-layer logits match model logits.", "kind": "method", "line": 254, "name": "test_fitted_late_layer_matches_model", "signature": "def test_fitted_late_layer_matches_model(self, fitted_lens, model)"}, {"doc": "Scenario: Explicit positions return correct subset.", "kind": "method", "line": 263, "name": "test_apply_with_explicit_positions", "signature": "def test_apply_with_explicit_positions(self, fitted_lens, model)"}, {"doc": "Scenario: use_jacobian=False returns untransported logits.", "kind": "method", "line": 274, "name": "test_logit_lens_baseline", "signature": "def test_logit_lens_baseline(self, fitted_lens, model)"}, {"doc": "Scenario: Unfitted layer raises ValueError.", "kind": "method", "line": 281, "name": "test_unfitted_layer_rejected", "signature": "def test_unfitted_layer_rejected(self, fitted_lens, model)"}, {"doc": "Scenario: Out-of-range layer raises ValueError.", "kind": "method", "line": 286, "name": "test_out_of_range_layer_rejected", "signature": "def test_out_of_range_layer_rejected(self, fitted_lens, model)"}, {"doc": "Scenario: merge() computes n_prompts-weighted mean.", "kind": "method", "line": 291, "name": "test_merge_weighted_mean", "signature": "def test_merge_weighted_mean(self)"}, {"doc": "Scenario: Mismatched lenses raise ValueError.", "kind": "method", "line": 319, "name": "test_merge_mismatch_raises", "signature": "def test_merge_mismatch_raises(self)"}, {"doc": "Scenario: Empty merge raises ValueError.", "kind": "method", "line": 326, "name": "test_merge_empty_raises", "signature": "def test_merge_empty_raises(self)"}, {"doc": "Scenario: transport() maps residual to final-layer basis.", "kind": "method", "line": 331, "name": "test_transport_produces_correct_shape", "signature": "def test_transport_produces_correct_shape(self, fitted_lens)"}, {"doc": "Scenario: Loading non-lens file raises ValueError.", "kind": "method", "line": 337, "name": "test_load_invalid_file_raises", "signature": "def test_load_invalid_file_raises(self, tmp_path)"}, {"doc": "Scenario: from_pretrained resolves a local file.", "kind": "method", "line": 344, "name": "test_from_pretrained_local_file", "signature": "def test_from_pretrained_local_file(self, fitted_lens, tmp_path)"}, {"doc": "Scenario: from_pretrained resolves a local directory.", "kind": "method", "line": 351, "name": "test_from_pretrained_local_directory", "signature": "def test_from_pretrained_local_directory(self, fitted_lens, tmp_path)"}, {"doc": "Scenario: repr contains key metadata.", "kind": "method", "line": 359, "name": "test_repr", "signature": "def test_repr(self, fitted_lens)"}, {"kind": "method", "line": 371, "name": "model", "signature": "def model(self)"}, {"doc": "Scenario: Resumed fit matches fresh fit.", "kind": "method", "line": 378, "name": "test_checkpoint_resume_produces_same_result", "signature": "def test_checkpoint_resume_produces_same_result(self, model, tmp_path)"}, {"doc": "Scenario: Resume after a skipped prompt does not double-count.\n\nRegression: a skipped prompt must not desync success-count from\nlist-position.", "kind": "method", "line": 408, "name": "test_resume_after_skip_no_double_count", "signature": "def test_resume_after_skip_no_double_count(self, model, tmp_path)"}, {"doc": "Scenario: Mismatched checkpoint settings raise ValueError.", "kind": "method", "line": 450, "name": "test_checkpoint_mismatch_raises", "signature": "def test_checkpoint_mismatch_raises(self, model, tmp_path)"}, {"doc": "Scenario: Default fit config has sensible defaults.", "kind": "method", "line": 476, "name": "test_fit_config_defaults", "signature": "def test_fit_config_defaults(self)"}, {"doc": "Scenario: Default app config has sensible defaults.", "kind": "method", "line": 485, "name": "test_app_config_defaults", "signature": "def test_app_config_defaults(self)"}, {"doc": "Scenario: Default app config uses all positions.", "kind": "method", "line": 497, "name": "test_default_config", "signature": "def test_default_config(self)"}, {"doc": "Scenario: Custom app config overrides specific layers.", "kind": "method", "line": 505, "name": "test_custom_config", "signature": "def test_custom_config(self)"}]}, {"id": "tests/test_lens_model.py", "kind": "module", "label": "test_lens_model.py", "language": "py", "sha256": "479f36e827cda9ce", "symbol_count": 34, "symbols": [{"doc": "Feature: TopoGPT3LensConfig provides centralized adapter configuration.", "kind": "class", "line": 13, "name": "TestTopoGPT3LensConfig", "signature": "class TestTopoGPT3LensConfig"}, {"doc": "Feature: TinyDecoder provides a minimal test model.", "kind": "class", "line": 41, "name": "TestTinyDecoder", "signature": "class TestTinyDecoder"}, {"doc": "Feature: TopoGPT3LensModel wraps a model to implement LensModel protocol.", "kind": "class", "line": 65, "name": "TestTopoGPT3LensModel", "signature": "class TestTopoGPT3LensModel"}, {"doc": "Feature: ActivationRecorder works with TopoGPT3LensModel.", "kind": "class", "line": 214, "name": "TestTopoGPT3LensModelWithRecording", "signature": "class TestTopoGPT3LensModelWithRecording"}, {"doc": "Feature: Edge cases are handled gracefully.", "kind": "class", "line": 278, "name": "TestTopoGPT3LensModelEdgeCases", "signature": "class TestTopoGPT3LensModelEdgeCases"}, {"doc": "Scenario: Default config matches small scale preset.", "kind": "method", "line": 16, "name": "test_default_config", "signature": "def test_default_config(self)"}, {"doc": "Scenario: Build lens config from TopoGPT2Config.", "kind": "method", "line": 25, "name": "test_from_topogpt2_config", "signature": "def test_from_topogpt2_config(self)"}, {"doc": "Scenario: Missing state.json raises FileNotFoundError.", "kind": "method", "line": 35, "name": "test_probe_checkpoint_missing_raises", "signature": "def test_probe_checkpoint_missing_raises(self, tmp_path)"}, {"doc": "Scenario: TinyDecoder has correct default shape.", "kind": "method", "line": 44, "name": "test_default_parameters", "signature": "def test_default_parameters(self)"}, {"doc": "Scenario: Forward pass produces correct logit shape.", "kind": "method", "line": 51, "name": "test_forward_output_shape", "signature": "def test_forward_output_shape(self)"}, {"doc": "Scenario: Embedding and LM head share weights.", "kind": "method", "line": 59, "name": "test_weight_tied", "signature": "def test_weight_tied(self)"}, {"kind": "method", "line": 69, "name": "raw_model", "signature": "def raw_model(self)"}, {"kind": "method", "line": 77, "name": "lens_model", "signature": "def lens_model(self, raw_model)"}, {"doc": "Scenario: LensModel attributes match underlying model.", "kind": "method", "line": 80, "name": "test_exposes_protocol_attributes", "signature": "def test_exposes_protocol_attributes(self, lens_model, raw_model)"}, {"doc": "Scenario: encode() returns tensor of shape [1, seq_len].", "kind": "method", "line": 87, "name": "test_encode_text_to_token_ids", "signature": "def test_encode_text_to_token_ids(self, lens_model)"}, {"doc": "Scenario: encode() uses BPETokenizer when available.", "kind": "method", "line": 95, "name": "test_encode_with_tokenizer", "signature": "def test_encode_with_tokenizer(self)"}, {"doc": "Scenario: encode() truncates at max_length.", "kind": "method", "line": 107, "name": "test_encode_respects_max_length", "signature": "def test_encode_respects_max_length(self, lens_model)"}, {"doc": "Scenario: forward() returns hidden states with d_model dim, not vocab.\n\nThe lens model forward should stop before final_norm and lm_head.\nThe output should have d_model as last dimension, not vocab_size.", "kind": "method", "line": 113, "name": "test_forward_returns_residual_only", "signature": "def test_forward_returns_residual_only(self)"}, {"doc": "Scenario: Residual forward shape differs from full model logits.", "kind": "method", "line": 128, "name": "test_forward_differs_from_full_model", "signature": "def test_forward_differs_from_full_model(self)"}, {"doc": "Scenario: unembed() maps residual to logits.", "kind": "method", "line": 141, "name": "test_unembed_produces_logits", "signature": "def test_unembed_produces_logits(self, lens_model)"}, {"doc": "Scenario: residual forward + unembed == model forward logits.\n\nThis validates that our split forward matches the original model's\nfull forward pass.", "kind": "method", "line": 150, "name": "test_forward_plus_unembed_matches_model_logits", "signature": "def test_forward_plus_unembed_matches_model_logits(self, lens_model, raw_model)"}, {"doc": "Scenario: Gradient flows through residual layers when grads enabled.", "kind": "method", "line": 163, "name": "test_autograd_graph_tracks_through_layers", "signature": "def test_autograd_graph_tracks_through_layers(self)"}, {"doc": "Scenario: input_device returns the embedding weight device.", "kind": "method", "line": 180, "name": "test_input_device_property", "signature": "def test_input_device_property(self, lens_model)"}, {"doc": "Scenario: input_device can be overridden.", "kind": "method", "line": 185, "name": "test_input_device_setter", "signature": "def test_input_device_setter(self, lens_model)"}, {"doc": "Scenario: tokenizer can be set after construction.", "kind": "method", "line": 191, "name": "test_tokenizer_setter", "signature": "def test_tokenizer_setter(self, lens_model)"}, {"doc": "Scenario: from_checkpoint with missing directory raises.", "kind": "method", "line": 198, "name": "test_from_checkpoint_missing_raises", "signature": "def test_from_checkpoint_missing_raises(self)"}, {"doc": "Scenario: Multiple forward passes with same input are deterministic.", "kind": "method", "line": 205, "name": "test_grad_enabled_deterministic", "signature": "def test_grad_enabled_deterministic(self, lens_model)"}, {"kind": "method", "line": 218, "name": "lens_model", "signature": "def lens_model(self)"}, {"doc": "Scenario: ActivationRecorder captures all requested layer outputs.", "kind": "method", "line": 225, "name": "test_recorder_captures_layer_outputs", "signature": "def test_recorder_captures_layer_outputs(self, lens_model)"}, {"doc": "Scenario: start_graph_at roots the autograd graph.", "kind": "method", "line": 238, "name": "test_recorder_with_start_graph_at", "signature": "def test_recorder_with_start_graph_at(self, lens_model)"}, {"doc": "Scenario: Hooks are removed even if construction fails.", "kind": "method", "line": 252, "name": "test_recorder_cleanup_on_exception", "signature": "def test_recorder_cleanup_on_exception(self, lens_model)"}, {"doc": "Scenario: Activations can be detached after recorder exits.", "kind": "method", "line": 264, "name": "test_recorder_detach_after_forward", "signature": "def test_recorder_detach_after_forward(self, lens_model)"}, {"doc": "Scenario: Empty input produces error or minimal output.", "kind": "method", "line": 281, "name": "test_empty_sequence", "signature": "def test_empty_sequence(self)"}, {"doc": "Scenario: Single token input works.", "kind": "method", "line": 291, "name": "test_single_token", "signature": "def test_single_token(self)"}]}, {"doc": "TopoGPT3: complex-valued spectral language model for code.  This package bundles:  - ``topogpt3.model``: the base TopoGPT2 architecture (quaternion spectral layers, BPE tokenizer, helpers). - ``topogpt3.train``: the curriculum trainer with Grassmannian / Fisher / phase diagnostics. - ``topogpt3.inference``: a standard autoregressive sampler that loads a trained safetensors checkpoint. - ``topogpt3.inference_hrm``: a hierarchical recursive reasoning sampler that reuses the same checkpoint with no extra trained parameters. - ``topogpt3.lens_model``: the Jacobian-lens model adapter (LensModel protocol + TopoGPT3LensModel wrapper). - ``topogpt3.jlens``: Jacobian lens fitting, application, and the ActivationRecorder / JacobianLens infrastructure.  Typical usage from a downstream project::  from topogpt3 import InferenceSettings, InferencePipeline  settings = InferenceSettings( checkpoint_dir=\"checkpoints_topogpt3\", prompt=\"def fibonacci(\", max_new_tokens=200, ) InferencePipeline(settings).execute()  Jacobian lens usage::", "id": "topogpt3/__init__.py", "kind": "module", "label": "__init__.py", "language": "py", "sha256": "f0c0693661707d14", "symbol_count": 0, "symbols": []}, {"id": "topogpt3/__main__.py", "kind": "module", "label": "__main__.py", "language": "py", "sha256": "1019f2b8207d812f", "symbol_count": 1, "symbols": [{"doc": "TopoGPT3 entry point. Delegates to subcommands.", "kind": "function", "line": 6, "name": "main", "signature": "def main()"}]}, {"doc": "OpenAI-compatible HTTP API server so TopoGPT3 can be used as a backend for coding agents (e.g. Pi, Aider, Continue, Codex CLI, etc.).  Security Posture ---------------- - **Authentication**: Bearer token (``Authorization: Bearer <key>``). Keys are loaded from ``--keys`` (comma-separated) or the ``TOPOGPT3_API_KEYS`` env var. Admin keys (prefixed ``admin:``) get higher rate limits. Constant-time comparison prevents timing leaks. - **Authorization**: token-bucket rate limiter per-key and per-IP with configurable thresholds. After ``max_failures`` bad auth attempts an IP is banned for ``ban_window`` seconds. - **Input hardening**: Pydantic schemas enforce strict types, min/max bounds, and length limits. Request body is capped server-side. Error responses never leak stack traces. - **Headers**: ``X-Content-Type-Options: nosniff``, ``X-Frame-Options: DENY``, ``X-XSS-Protection: 1; mode=block``, ``Content-Security-Policy: default-src 'none'`` on every response. CORS policy allows nothing by default (configurable allow-origins). - **Audit**: structured JSON log lines for every request (truncated bodies, no secrets).  Usage::  TOPOGPT3_API_KEYS=\"sk-secret-key,admin:sk-admin-key\" \\\\ python -m topogpt3 api_server \\\\ --checkpoint checkpoints_topogpt3/last \\\\ --port 8800  Pi / agent config::", "id": "topogpt3/api_server.py", "kind": "module", "label": "api_server.py", "language": "py", "sha256": "0bc31521f6b07f4c", "symbol_count": 46, "symbols": [{"kind": "function", "line": 116, "name": "_setup_logging", "signature": "def _setup_logging(verbose)"}, {"kind": "class", "line": 137, "name": "ApiKey", "signature": "class ApiKey"}, {"kind": "class", "line": 143, "name": "AuthState", "signature": "class AuthState"}, {"doc": "Accept ``key1,admin:key2,key3``. The ``admin:`` prefix marks an\nadmin-level key; everything else is a regular user key.", "kind": "method", "line": 164, "name": "_parse_keys", "signature": "def _parse_keys(raw)"}, {"kind": "method", "line": 192, "name": "_sha256", "signature": "def _sha256(raw)"}, {"kind": "class", "line": 202, "name": "TokenBucket", "signature": "class TokenBucket"}, {"kind": "class", "line": 219, "name": "RateLimiter", "signature": "class RateLimiter"}, {"kind": "class", "line": 250, "name": "IpBanner", "signature": "class IpBanner"}, {"kind": "method", "line": 281, "name": "_sanitize_stop", "signature": "def _sanitize_stop(stop)"}, {"kind": "class", "line": 291, "name": "CompletionRequest", "signature": "class CompletionRequest(BaseModel)"}, {"kind": "class", "line": 310, "name": "Message", "signature": "class Message(BaseModel)"}, {"kind": "class", "line": 316, "name": "ChatCompletionRequest", "signature": "class ChatCompletionRequest(BaseModel)"}, {"kind": "class", "line": 341, "name": "ServerModel", "signature": "class ServerModel"}, {"kind": "method", "line": 499, "name": "_resolve_device", "signature": "def _resolve_device(device)"}, {"doc": "Probe a checkpoint and return (scale, n_kv_heads).\n\nReads the model hidden dim from ``final_norm.weight`` (or the first\ntransformer layer's output projection) and the KV head count from the\n``k_proj`` output dim, so the server can load checkpoints of any scale\n(small/large/…) without hardcoding the preset.", "kind": "method", "line": 505, "name": "_probe_arch", "signature": "def _probe_arch(checkpoint_dir)"}, {"kind": "method", "line": 553, "name": "load_model", "signature": "def load_model(checkpoint, device)"}, {"kind": "method", "line": 577, "name": "lifespan", "signature": "def lifespan(app)"}, {"doc": "Global middleware: rate-limit, IP-ban, security headers, audit log.", "kind": "method", "line": 617, "name": "_security_middleware", "signature": "def _security_middleware(request, call_next)"}, {"doc": "Best-effort real client IP. We trust no proxy headers by default.", "kind": "method", "line": 645, "name": "_real_ip", "signature": "def _real_ip(request)"}, {"kind": "method", "line": 656, "name": "_json_error", "signature": "def _json_error(status, detail)"}, {"doc": "FastAPI dependency: extract & validate Bearer token.\n\nWhen no API keys are configured (auth disabled via ``--no-auth`` or an\nempty ``TOPOGPT3_API_KEYS``), the dependency is a pass-through so local\nintegrations (e.g. the ExploitGym TopoExploit agent) can call it without\na token.", "kind": "method", "line": 668, "name": "_authenticate", "signature": "def _authenticate(request)"}, {"doc": "Rate limit per-key (with admin exemption / higher limit).", "kind": "method", "line": 694, "name": "_check_rate_limit", "signature": "def _check_rate_limit(api_key, request)"}, {"kind": "method", "line": 712, "name": "health", "signature": "def health(request)"}, {"kind": "method", "line": 719, "name": "list_models", "signature": "def list_models(request)"}, {"kind": "method", "line": 736, "name": "completions", "signature": "def completions(req, request)"}, {"kind": "method", "line": 792, "name": "chat_completions", "signature": "def chat_completions(req, request)"}, {"kind": "method", "line": 850, "name": "_check_model", "signature": "def _check_model()"}, {"kind": "method", "line": 855, "name": "_short_id", "signature": "def _short_id()"}, {"kind": "method", "line": 859, "name": "_build_chat_prompt", "signature": "def _build_chat_prompt(messages)"}, {"kind": "method", "line": 866, "name": "_extract_text", "signature": "def _extract_text(content)"}, {"kind": "method", "line": 880, "name": "_stream_completion", "signature": "def _stream_completion(prompt, max_tokens, temperature, top_k, repetition_penalty, stop, auto_continue, max_continuations)"}, {"kind": "method", "line": 915, "name": "_stream_chat", "signature": "def _stream_chat(t0_ms, prompt, max_tokens, temperature, top_k, repetition_penalty, stop, auto_continue, max_continuations)"}, {"kind": "method", "line": 953, "name": "main", "signature": "def main()"}, {"kind": "method", "line": 148, "name": "validate", "signature": "def validate(self, raw)"}, {"kind": "method", "line": 208, "name": "consume", "signature": "def consume(self, n)"}, {"kind": "method", "line": 220, "name": "__init__", "signature": "def __init__(self, user_rps, admin_rps, capacity)"}, {"kind": "method", "line": 227, "name": "_cleanup", "signature": "def _cleanup(self)"}, {"kind": "method", "line": 233, "name": "allow", "signature": "def allow(self, key, role)"}, {"kind": "method", "line": 251, "name": "__init__", "signature": "def __init__(self, max_failures, window)"}, {"kind": "method", "line": 257, "name": "record_failure", "signature": "def record_failure(self, ip)"}, {"kind": "method", "line": 265, "name": "is_banned", "signature": "def is_banned(self, ip)"}, {"kind": "method", "line": 306, "name": "_normalize_stop", "signature": "def _normalize_stop(cls, v)"}, {"kind": "method", "line": 331, "name": "_normalize_stop", "signature": "def _normalize_stop(cls, v)"}, {"kind": "method", "line": 348, "name": "complete", "signature": "def complete(self, prompt)"}, {"kind": "method", "line": 393, "name": "stream_complete", "signature": "def stream_complete(self, prompt)"}, {"kind": "method", "line": 482, "name": "_is_eos", "signature": "def _is_eos(self, token_id)"}]}, {"doc": "Auto-continuation engine: detects truncated responses and feeds the last incomplete lines back so the model can resume where it left off.  Used by both the standard inference pipeline and the HRM \"thinking\" mode.", "id": "topogpt3/continuation.py", "kind": "module", "label": "continuation.py", "language": "py", "sha256": "9d0c4e1576ec926a", "symbol_count": 5, "symbols": [{"kind": "function", "line": 25, "name": "_count_unclosed_brackets", "signature": "def _count_unclosed_brackets(text)"}, {"kind": "function", "line": 36, "name": "_count_unclosed_fences", "signature": "def _count_unclosed_fences(text)"}, {"doc": "Heuristic to decide whether a model response looks finished.\n\nReturns True when the response seems naturally complete (no need to\ncontinue), False when it appears truncated and continuation may help.", "kind": "function", "line": 45, "name": "is_response_complete", "signature": "def is_response_complete(text, min_chars)"}, {"doc": "Return the last N lines (or up to tail_chars) of `text` as a\ncontinuation prefix to feed back into the model.\n\nThe returned string can be prepended as context for the model's next\ngeneration call so it continues naturally from that point.", "kind": "function", "line": 75, "name": "extract_tail_for_continuation", "signature": "def extract_tail_for_continuation(text, tail_lines, tail_chars)"}, {"doc": "Split `text` at the last newline.\n\nReturns (prefix_without_last_line, last_line).\nUseful for discarding a trailing incomplete line before continuation.", "kind": "function", "line": 105, "name": "split_at_last_newline", "signature": "def split_at_last_newline(text)"}]}, {"doc": "Elastic Weight Consolidation (EWC) and Experience Replay for TopoGPT3.  Provides EWCRegularizer for regularizing against catastrophic forgetting across curriculum tiers, and ReplayBuffer for interleaving data from previous tiers during training.  Author: Gris Iscomeback License: GPL v3", "id": "topogpt3/ewc.py", "kind": "module", "label": "ewc.py", "language": "py", "sha256": "7c1d412f69d3a374", "symbol_count": 11, "symbols": [{"doc": "Elastic Weight Consolidation regularizer.\n\nComputes the diagonal Fisher information matrix over a dataset and uses\nit to penalize large deviations from the parameters that were optimal\nfor previous tasks (tiers). The penalty is:\n\n    L_EWC = lambda * sum_i  F_i * (theta_i - theta_opt_i)^2\n\nwhere F_i is the Fisher information for parameter i and theta_opt_i\nis the parameter value at the end of the previous task.\n\nUsage:\n    ewc = EWCRegularizer(model, lambda_ewc=5000.0)\n    ewc.compute_fisher(dataloader, vocab_size, device)\n    # ... train on new tier ...\n    loss = task_loss + ewc.penalty(model)", "kind": "class", "line": 25, "name": "EWCRegularizer", "signature": "class EWCRegularizer"}, {"doc": "Experience Replay buffer for interleaving data from previous curriculum\ntiers during training.\n\nMaintains references to the training dataloaders of previously completed\ntiers and mixes their batches into the current tier's training loop at a\nconfigurable ratio.\n\nUsage:\n    replay = ReplayBuffer(max_tiers=3, replay_ratio=0.25)\n    replay.add_tier(0, tier0_train_dl)\n    replay.add_tier(1, tier1_train_dl)\n    # During tier-2 training, for each step:\n    bx, by = replay.sample_batch(current_batch_size=4)\n    # bx, by contains replay data from tiers 0,1", "kind": "class", "line": 223, "name": "ReplayBuffer", "signature": "class ReplayBuffer"}, {"doc": "Args:\n    model: The model to regularize.\n    lambda_ewc: Penalty strength scaling the Fisher-weighted\n        parameter deviation loss.\n    fisher_num_samples: Number of batches used when computing the\n        Fisher information matrix (stored for reference).", "kind": "method", "line": 45, "name": "__init__", "signature": "def __init__(self, model, lambda_ewc, fisher_num_samples)"}, {"doc": "Compute the diagonal Fisher information matrix and snapshot the\ncurrent parameters as the reference point for the EWC penalty.\n\nFor each batch: forward pass -> per-sample NLL -> backprop ->\naccumulate squared gradients.  Handles AMP autocast.\n\nAfter calling this method, ``self.fisher`` and ``self.old_params``\nare populated and ready for ``penalty()``.\n\nArgs:\n    dataloader: Iterable yielding ``(bx, by)`` batches.\n    vocab_size: Size of the vocabulary (for reshaping logits).\n    device: Device where the model lives.\n    max_batches: Maximum number of batches to use for estimation.", "kind": "method", "line": 68, "name": "compute_fisher", "signature": "def compute_fisher(self, dataloader, vocab_size, device, max_batches)"}, {"doc": "Compute the EWC regularization loss.\n\nReturns:\n    Scalar tensor: lambda * sum_i F_i * (theta_i - theta_opt_i)^2", "kind": "method", "line": 168, "name": "penalty", "signature": "def penalty(self, model)"}, {"doc": "Persist the Fisher information matrix and old parameter snapshots\nto a single ``.pt`` file.\n\nArgs:\n    path: Destination file path (should end in ``.pt``).", "kind": "method", "line": 186, "name": "save_state", "signature": "def save_state(self, path)"}, {"doc": "Load a previously saved EWC state.\n\nArgs:\n    path: Path to the ``.pt`` file created by ``save_state``.", "kind": "method", "line": 203, "name": "load_state", "signature": "def load_state(self, path)"}, {"doc": "Args:\n    max_tiers: Maximum number of previous tiers to remember.\n        When exceeded, the oldest tier is dropped.\n    replay_ratio: Fraction of the returned batch that comes from\n        replayed tiers (0.0 = no replay, 1.0 = all replay).", "kind": "method", "line": 241, "name": "__init__", "signature": "def __init__(self, max_tiers, replay_ratio)"}, {"doc": "Register a tier's training dataloader for future replay.\n\nIf the buffer is full (``max_tiers`` tiers stored), the oldest\nentry is evicted.\n\nArgs:\n    tier_index: Integer tier identifier (e.g. 0, 1, 2, ...).\n    dataloader: The training dataloader for that tier.", "kind": "method", "line": 259, "name": "add_tier", "signature": "def add_tier(self, tier_index, dataloader)"}, {"doc": "Get the next batch from a tier's dataloader, cycling if necessary.\n\nArgs:\n    tier_index: The tier identifier.\n    dataloader: The tier's dataloader.\n\nReturns:\n    ``(bx, by)`` tensors or ``None`` if the dataloader is empty.", "kind": "method", "line": 282, "name": "_get_next_batch", "signature": "def _get_next_batch(self, tier_index, dataloader)"}, {"doc": "Sample a mixed replay batch from ALL stored previous tiers.\n\nThe returned batch has size approximately\n``replay_ratio * current_batch_size`` and is drawn uniformly from\nall stored tiers, cycling through their dataloaders.\n\nThe caller is responsible for concatenating this with the current\ntier's batch::\n\n    current_batch = next(current_dataloader)\n    replay = replay_buffer.sample_batch(current_batch_size)\n    if replay is not None:\n        bx = torch.cat([current_batch[0], replay[0]], dim=0)\n        by = torch.cat([current_batch[1], replay[1]], dim=0)\n\nArgs:\n    current_batch_size: The batch size of the current tier.\n        Used to compute the target replay batch size as\n        ``max(1, int(replay_ratio * current_batch_size))``.\n\nReturns:\n    ``(bx, by)`` tensors on the same device as the dataloader\n    yields them, or ``None`` if no replay tiers are available.", "kind": "method", "line": 312, "name": "sample_batch", "signature": "def sample_batch(self, current_batch_size)"}]}, {"doc": "TopoExploit Training Configuration  122M parameter model trained on ExploitGym vulnerability/exploit data.  Tier structure: Tier 1: vulnerability_analysis (256 seq, 3 epochs) Tier 2: patch_analysis        (512 seq, 4 epochs) Tier 3: exploit_development   (1024 seq, 5 epochs)  Author: Gris Iscomeback License: GPL v3", "id": "topogpt3/exploitgym_config.py", "kind": "module", "label": "exploitgym_config.py", "language": "py", "sha256": "404cea85d4c207fa", "symbol_count": 3, "symbols": [{"doc": "Configuration for TopoExploit: 122M params on ExploitGym data.", "kind": "class", "line": 26, "name": "TopoExploitConfig", "signature": "class TopoExploitConfig"}, {"kind": "method", "line": 100, "name": "build_topogpt2_config", "signature": "def build_topogpt2_config(self, max_seq_len, attn_window)"}, {"kind": "method", "line": 122, "name": "build_loader_config", "signature": "def build_loader_config(self)"}]}, {"doc": "ExploitGym Data Loader for TopoExploit  Clones the ExploitGym repo and tokenizes task data into training memmaps. Each task contains vulnerability descriptions, patches, exploit docs, and PoC source.  Tier structure: Tier 1: vulnerability_analysis -- CVE description -> vulnerability analysis Tier 2: patch_analysis        -- source + patch -> fix explanation Tier 3: exploit_development   -- vulnerability -> exploit strategy + PoC code  Author: Gris Iscomeback License: GPL v3", "id": "topogpt3/exploitgym_loader.py", "kind": "module", "label": "exploitgym_loader.py", "language": "py", "sha256": "8c9e2a44c0355f8e", "symbol_count": 19, "symbols": [{"kind": "class", "line": 39, "name": "ExploitGymLoaderConfig", "signature": "class ExploitGymLoaderConfig"}, {"kind": "method", "line": 54, "name": "ensure_repo", "signature": "def ensure_repo(repo_cache, logger)"}, {"kind": "method", "line": 75, "name": "_read_file_safe", "signature": "def _read_file_safe(path, max_chars)"}, {"kind": "method", "line": 83, "name": "_parse_task_ids", "signature": "def _parse_task_ids(repo_path)"}, {"kind": "method", "line": 93, "name": "_task_id_to_path", "signature": "def _task_id_to_path(task_id)"}, {"kind": "method", "line": 101, "name": "parse_task", "signature": "def parse_task(task_dir, task_id, max_chars)"}, {"kind": "method", "line": 137, "name": "_format_vulnerability_analysis", "signature": "def _format_vulnerability_analysis(task)"}, {"kind": "method", "line": 153, "name": "_format_patch_analysis", "signature": "def _format_patch_analysis(task)"}, {"kind": "method", "line": 168, "name": "_format_exploit_development", "signature": "def _format_exploit_development(task)"}, {"kind": "class", "line": 200, "name": "ExploitGymDataLoader", "signature": "class ExploitGymDataLoader"}, {"kind": "method", "line": 201, "name": "__init__", "signature": "def __init__(self, config, tokenizer, logger)"}, {"kind": "method", "line": 209, "name": "_tier_paths", "signature": "def _tier_paths(self, tier)"}, {"kind": "method", "line": 215, "name": "_manifest_path", "signature": "def _manifest_path(self, tier)"}, {"kind": "method", "line": 218, "name": "_already_prepared", "signature": "def _already_prepared(self, tier)"}, {"kind": "method", "line": 240, "name": "_load_all_tasks", "signature": "def _load_all_tasks(self, repo_path)"}, {"kind": "method", "line": 264, "name": "prepare_tier", "signature": "def prepare_tier(self, tier_index, force)"}, {"kind": "method", "line": 411, "name": "open_memmap", "signature": "def open_memmap(self, tier, split)"}, {"kind": "method", "line": 288, "name": "stratum_key", "signature": "def stratum_key(task)"}, {"kind": "method", "line": 343, "name": "flush", "signature": "def flush(split)"}]}, {"doc": "Fases 2-5: martillo Hodge-CM para TopoGPT3 / modelo espectral continuo. F2: CM phase-lattice STE suave. F3: Hodge regularizer + Fisher hinge. F4: MCMC offline. F5: Hecke pooling solo escala.", "id": "topogpt3/hodge_cm.py", "kind": "module", "label": "hodge_cm.py", "language": "py", "sha256": "dcd4b143bbee18ac", "symbol_count": 15, "symbols": [{"kind": "class", "line": 13, "name": "CMPhaseQuantSTE", "signature": "class CMPhaseQuantSTE(Function)"}, {"doc": "E_Q: distancia cuadratica media a raices m-esimas. kr,ki reales misma forma.", "kind": "method", "line": 25, "name": "cm_phase_loss", "signature": "def cm_phase_loss(kr, ki, m)"}, {"doc": "Cuantizacion suave de fase con STE. Magnitud intacta.", "kind": "method", "line": 34, "name": "apply_cm_soft", "signature": "def apply_cm_soft(kr, ki, m, beta)"}, {"kind": "method", "line": 49, "name": "freq_grid_laplacian", "signature": "def freq_grid_laplacian(fh, fw, device, dtype)"}, {"doc": "||K - e^{-tL} K||_F^2 sobre grid freq. kr,ki: [...,Fh,Fw].", "kind": "method", "line": 66, "name": "hodge_heat_residue", "signature": "def hodge_heat_residue(kr, ki, t)"}, {"kind": "method", "line": 81, "name": "harmonic_ratio", "signature": "def harmonic_ratio(kr, ki, t)"}, {"kind": "method", "line": 96, "name": "fisher_hinge", "signature": "def fisher_hinge(singulars, r, margin)"}, {"doc": "ULA offline: theta <- theta - eta grad E + sqrt(2 eta T) xi.", "kind": "method", "line": 105, "name": "langevin_refine", "signature": "def langevin_refine(param, energy_fn, steps, eta, temp, seed)"}, {"doc": "Clustering blando por fase para nodos/frecuencias. feats: [N,D] complejo o real.", "kind": "method", "line": 123, "name": "gibbs_cluster_assign", "signature": "def gibbs_cluster_assign(feats, n_clusters, iters, seed)"}, {"doc": "Promedio sobre orbitas de escala/traslacion. Solo donde hay simetria real\n(Strassen T, Fourier interp Willmore). NO reemplazo de atencion general.", "kind": "class", "line": 147, "name": "HeckeScalePool", "signature": "class HeckeScalePool(Module)"}, {"kind": "method", "line": 15, "name": "forward", "signature": "def forward(ctx, phase, m, beta)"}, {"kind": "method", "line": 21, "name": "backward", "signature": "def backward(ctx, grad_out)"}, {"kind": "method", "line": 53, "name": "idx", "signature": "def idx(h, w)"}, {"kind": "method", "line": 151, "name": "__init__", "signature": "def __init__(self, modes)"}, {"kind": "method", "line": 155, "name": "forward", "signature": "def forward(self, Kfreq)"}]}, {"doc": "TopoGPT3 inference engine.  Production-grade autoregressive code completion pipeline for TopoGPT3 checkpoints. Loads weights from safetensors, aligns the underlying TopoGPT2 architecture against the stored tensors, optionally applies the Gauss complex-multiply patch for numerical parity with training, and performs sampling with repetition penalty and top-k filtering.  The pipeline is decomposed into single-responsibility collaborators wired by an orchestrator. All paths, sampling parameters, safety bounds and string identifiers live inside InferenceSettings so that business logic contains no magic numbers or hardcoded constants.", "id": "topogpt3/inference.py", "kind": "module", "label": "inference.py", "language": "py", "sha256": "ab6c777c0bf86c8c", "symbol_count": 54, "symbols": [{"doc": "Immutable architecture preset for a named model scale.", "kind": "class", "line": 32, "name": "ScalePreset", "signature": "class ScalePreset"}, {"doc": "Centralized configuration container for the inference pipeline.\n\nEvery value consumed downstream resides here. Adding a new tunable means\nextending this class; no other module should embed literals.", "kind": "class", "line": 42, "name": "InferenceSettings", "signature": "class InferenceSettings"}, {"doc": "Builds a stdout-attached logger from inference settings.", "kind": "class", "line": 160, "name": "InferenceLoggerFactory", "signature": "class InferenceLoggerFactory"}, {"doc": "Resolves filesystem paths while rejecting traversal outside their root.", "kind": "class", "line": 180, "name": "SecurePathResolver", "signature": "class SecurePathResolver"}, {"doc": "Resolves the TopoGPT3 runtime module via the package import system.", "kind": "class", "line": 214, "name": "SourceModuleLoader", "signature": "class SourceModuleLoader"}, {"doc": "Computes and validates checkpoint file paths under a single root.", "kind": "class", "line": 230, "name": "CheckpointPaths", "signature": "class CheckpointPaths"}, {"doc": "Reads tensor metadata from safetensors to infer architecture details.", "kind": "class", "line": 276, "name": "WeightShapeProbe", "signature": "class WeightShapeProbe"}, {"doc": "Builds a TopoGPT2Config matching the loaded checkpoint and tokenizer.", "kind": "class", "line": 320, "name": "TopoGPT2ConfigAligner", "signature": "class TopoGPT2ConfigAligner"}, {"doc": "Builds a BPETokenizer instance using the configured encoding.", "kind": "class", "line": 351, "name": "TokenizerFactory", "signature": "class TokenizerFactory"}, {"doc": "Applies the idempotent Gauss complex-multiply patch when enabled.", "kind": "class", "line": 364, "name": "GaussPatchApplier", "signature": "class GaussPatchApplier"}, {"doc": "Instantiates the model and loads weights from safetensors.", "kind": "class", "line": 382, "name": "ModelAssembler", "signature": "class ModelAssembler"}, {"doc": "Applies deterministic seeds across torch, CUDA and the model package.", "kind": "class", "line": 419, "name": "SeedSynchronizer", "signature": "class SeedSynchronizer"}, {"doc": "Immutable sampling parameters consumed by the generation engine.", "kind": "class", "line": 443, "name": "SamplingPolicy", "signature": "class SamplingPolicy"}, {"doc": "Quantitative summary of a single generation call.", "kind": "class", "line": 463, "name": "GenerationReport", "signature": "class GenerationReport"}, {"doc": "Runs autoregressive sampling against a loaded model and tokenizer.", "kind": "class", "line": 477, "name": "GenerationEngine", "signature": "class GenerationEngine"}, {"doc": "Prints a GenerationReport to stdout using settings-defined formatting.", "kind": "class", "line": 535, "name": "ResultRenderer", "signature": "class ResultRenderer"}, {"doc": "Orchestrator wiring loader, builder, engine and renderer.", "kind": "class", "line": 564, "name": "InferencePipeline", "signature": "class InferencePipeline"}, {"doc": "Translates command-line arguments into an InferenceSettings instance.", "kind": "class", "line": 617, "name": "CliArgumentParser", "signature": "class CliArgumentParser"}, {"doc": "CLI entry point. Returns a process exit code.", "kind": "method", "line": 723, "name": "main", "signature": "def main(argv)"}, {"doc": "Return the architecture preset table indexed by scale name.", "kind": "method", "line": 103, "name": "scale_presets", "signature": "def scale_presets()"}, {"doc": "Return the resolved preset for the configured model scale.", "kind": "method", "line": 118, "name": "preset", "signature": "def preset(self)"}, {"doc": "Raise ValueError if any setting falls outside its safety bounds.", "kind": "method", "line": 128, "name": "validate", "signature": "def validate(self)"}, {"doc": "Return a configured Logger with a single deduplicated stdout handler.", "kind": "method", "line": 164, "name": "build", "signature": "def build(settings)"}, {"doc": "Join `parts` under `root` and return the canonical resolved path.\n\nRaises ValueError if the resolved path escapes `root`.", "kind": "method", "line": 184, "name": "resolve_under", "signature": "def resolve_under(root)"}, {"doc": "Validate `path` points to an existing regular file with the expected suffix.", "kind": "method", "line": 200, "name": "require_existing_file", "signature": "def require_existing_file(path, expected_suffix)"}, {"kind": "method", "line": 217, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Return the topogpt3.train module which re-exports model symbols.", "kind": "method", "line": 221, "name": "load", "signature": "def load(self)"}, {"kind": "method", "line": 233, "name": "__init__", "signature": "def __init__(self, settings)"}, {"doc": "Directory holding the active checkpoint slot.", "kind": "method", "line": 241, "name": "slot_dir", "signature": "def slot_dir(self)"}, {"doc": "Resolved path to the safetensors weights file inside the slot.", "kind": "method", "line": 245, "name": "model_file", "signature": "def model_file(self)"}, {"doc": "Resolved path to the JSON training-state file inside the slot.", "kind": "method", "line": 251, "name": "state_file", "signature": "def state_file(self)"}, {"doc": "Verify weights exist and the on-disk size lies within safety bounds.", "kind": "method", "line": 257, "name": "assert_ready", "signature": "def assert_ready(self)"}, {"kind": "method", "line": 279, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Recover N_KV_HEADS used at training by inspecting the k_proj shape.\n\nReturns None when the probe key is absent, signalling the caller to\nfall back to scale defaults rather than guess.", "kind": "method", "line": 283, "name": "detect_n_kv_heads", "signature": "def detect_n_kv_heads(self, weights_path, d_model, n_heads)"}, {"kind": "method", "line": 323, "name": "__init__", "signature": "def __init__(self, settings, source_module, logger)"}, {"doc": "Return a TopoGPT2Config dataclass ready to instantiate the model.", "kind": "method", "line": 329, "name": "build", "signature": "def build(self, n_kv_heads, vocab_size)"}, {"kind": "method", "line": 354, "name": "__init__", "signature": "def __init__(self, settings, source_module)"}, {"doc": "Return an instance of BPETokenizer bound to the configured encoding.", "kind": "method", "line": 358, "name": "build", "signature": "def build(self)"}, {"kind": "method", "line": 367, "name": "__init__", "signature": "def __init__(self, settings, source_module, logger)"}, {"doc": "Patch QuaternionSpectralLayer to use the 3-multiply Gauss contract.", "kind": "method", "line": 373, "name": "apply_if_enabled", "signature": "def apply_if_enabled(self)"}, {"kind": "method", "line": 385, "name": "__init__", "signature": "def __init__(self, settings, source_module, logger)"}, {"doc": "Build the TopoGPT2 graph, load weights into it, and return it in eval mode.", "kind": "method", "line": 391, "name": "assemble", "signature": "def assemble(self, aligned_cfg, paths)"}, {"kind": "method", "line": 422, "name": "__init__", "signature": "def __init__(self, settings, source_module, logger)"}, {"doc": "Seed all relevant RNGs using the model package helper when available.", "kind": "method", "line": 428, "name": "apply", "signature": "def apply(self)"}, {"doc": "Construct a SamplingPolicy from inference settings.", "kind": "method", "line": 452, "name": "from_settings", "signature": "def from_settings(cls, settings)"}, {"doc": "Return throughput in tokens/sec, clamped to avoid divide-by-zero.", "kind": "method", "line": 472, "name": "tokens_per_second", "signature": "def tokens_per_second(self, elapsed_floor)"}, {"kind": "method", "line": 480, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Generate a completion for `prompt` and return a GenerationReport.", "kind": "method", "line": 485, "name": "run", "signature": "def run(self, model, tokenizer, prompt, policy)"}, {"kind": "method", "line": 538, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Emit a banner with prompt and completion, plus a throughput log line.", "kind": "method", "line": 542, "name": "render", "signature": "def render(self, report)"}, {"kind": "method", "line": 567, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Run the full inference pipeline end-to-end and return the report.", "kind": "method", "line": 573, "name": "execute", "signature": "def execute(self)"}, {"doc": "Return the configured argparse.ArgumentParser.", "kind": "method", "line": 621, "name": "build_parser", "signature": "def build_parser()"}, {"doc": "Parse `argv` (or sys.argv) and return a populated InferenceSettings.", "kind": "method", "line": 700, "name": "parse", "signature": "def parse(argv)"}]}, {"doc": "TopoGPT3.1: Hierarchical Recursive Reasoning Inference Engine.  This module extends TopoGPT3 with a parameter-free hierarchical recursive reasoning pipeline inspired by:  * Hierarchical Reasoning Model (HRM), Sapient Intelligence: a biologically motivated two-speed architecture with a slow high-level loop and a fast low-level loop. * Tiny Recursive Model (TRM) and Generative Recursive Reasoning Models (GRAM): latent-space recurrence that iterates token vectors until they reach an attractor before projecting them outward.  The pipeline is intentionally built so that the underlying TopoGPT2 weight matrices remain bit-identical to those produced by the TopoGPT3 trainer. No new learnable parameters are introduced. The pretrained transformer layers are repurposed as the recurrent step function of a hierarchical fixed-point iteration whose halting condition is the empirical stabilization of the latent state.  The high-level slow state is persisted across multiple emitted tokens to achieve sparse temporal reasoning: the full network is iterated only at configurable intervals, while a short suffix of layers refines the low-level state at every emitted token.  All configurable values reside in dedicated configuration dataclasses; no magic numbers or hardcoded constants are embedded in business logic. Path resolution rejects traversal escapes. State dict loading defers strictness to settings, so an architecturally aligned TopoGPT3 checkpoint loads unchanged.", "id": "topogpt3/inference_hrm.py", "kind": "module", "label": "inference_hrm.py", "language": "py", "sha256": "a037ba082478370d", "symbol_count": 76, "symbols": [{"doc": "Immutable architecture preset for a named model scale.", "kind": "class", "line": 54, "name": "ScalePreset", "signature": "class ScalePreset"}, {"doc": "Hyperparameters governing the hierarchical recursive thinking loop.\n\nThe semantics follow the HRM and GRAM literature, adapted to operate\nsafely with zero additional learnable parameters on a model that was\nnot trained with recurrence in its computational graph. The reasoner\nperforms damped fixed-point iteration entirely in the residual-stream\nspace produced by the baseline forward pass; deep activations are never\nfed back into the token-embedding-input layers, preserving the trained\nactivation distribution at every layer boundary.\n\nAttributes:\n    enabled: master switch; when False the pipeline degrades to the\n        standard non-recursive autoregressive loop.\n    max_high_level_iters: maximum slow-loop iterations per emitted token.\n        Each iteration applies a deeper trailing window of layers.\n    max_low_level_iters: maximum fast-loop iterations per high-level step.\n        Each iteration applies the short trailing window of layers.\n    low_level_window: number of trailing transformer layers iterated by\n        the low-level fast loop.\n    high_level_window: number of trailing transformer layers iterated by\n        the high-level slow loop. Should be greater than or equal to\n        low_level_window so the hierarchy matches the HRM coarse/fine\n        split.\n    low_level_step: damping coefficient in [0, 1] for the low-level\n        update rule z <- z + step * (window(z) - z).\n    high_level_step: damping coefficient for the high-level update.\n    attractor_low_epsilon: relative L2 change threshold that declares the\n        low-level state converged.\n    attractor_high_epsilon: relative L2 change threshold that declares the\n        high-level state converged.\n    high_level_persist_tokens: tokens during which the refinement vector\n        is reused as a warm start before being re-initialized to zero.\n        This is the sparse temporal-memory dimension.\n    cache_warm_start_weight: scalar in [0, 1] applied to the cached\n        refinement before warm-starting the next token's iteration.\n    max_drift_relative: relative L2 distance ceiling between the iterated\n        latent and the baseline latent; exceeding it triggers a reset to\n        the baseline state and aborts thinking for the current token.\n    latent_change_eps: floor used in the denominator of relative change\n        computations to avoid division by zero.\n    safety_max_total_iterations: hard cap on total layer invocations per\n        emitted token regardless of configured iters.\n    minimum_low_level_iters: floor on low-level iterations before\n        convergence checks may halt the loop.\n    minimum_high_level_iters: floor on high-level iterations before\n        convergence checks may halt the loop.\n    diagnostic_logging: when True, emits per-token iteration statistics.", "kind": "class", "line": 64, "name": "RecursiveReasoningConfig", "signature": "class RecursiveReasoningConfig"}, {"doc": "Centralized configuration for the TopoGPT3.1 inference pipeline.\n\nEvery value consumed downstream resides here. Extending the pipeline with\na new tunable means extending this dataclass; no other module should\nembed literals.", "kind": "class", "line": 134, "name": "HRMInferenceSettings", "signature": "class HRMInferenceSettings"}, {"doc": "Builds a stdout-attached logger from inference settings.", "kind": "class", "line": 343, "name": "HRMLoggerFactory", "signature": "class HRMLoggerFactory"}, {"doc": "Resolves filesystem paths while rejecting traversal outside their root.", "kind": "class", "line": 363, "name": "SecurePathResolver", "signature": "class SecurePathResolver"}, {"doc": "Resolves the TopoGPT3 runtime module via the package import system.", "kind": "class", "line": 397, "name": "SourceModuleLoader", "signature": "class SourceModuleLoader"}, {"doc": "Computes and validates checkpoint file paths under a single root.", "kind": "class", "line": 413, "name": "CheckpointPaths", "signature": "class CheckpointPaths"}, {"doc": "Reads tensor metadata from safetensors to infer architecture details.", "kind": "class", "line": 459, "name": "WeightShapeProbe", "signature": "class WeightShapeProbe"}, {"doc": "Builds a TopoGPT2Config matching the loaded checkpoint and tokenizer.", "kind": "class", "line": 502, "name": "TopoGPT2ConfigAligner", "signature": "class TopoGPT2ConfigAligner"}, {"doc": "Builds a BPETokenizer instance using the configured encoding.", "kind": "class", "line": 533, "name": "TokenizerFactory", "signature": "class TokenizerFactory"}, {"doc": "Applies the idempotent Gauss complex-multiply patch when enabled.", "kind": "class", "line": 546, "name": "GaussPatchApplier", "signature": "class GaussPatchApplier"}, {"doc": "Instantiates the model and loads weights from safetensors.", "kind": "class", "line": 564, "name": "ModelAssembler", "signature": "class ModelAssembler"}, {"doc": "Applies deterministic seeds across torch, CUDA and the model package.", "kind": "class", "line": 601, "name": "SeedSynchronizer", "signature": "class SeedSynchronizer"}, {"doc": "Computes the relative L2 distance between two latent tensors.", "kind": "class", "line": 624, "name": "LatentChangeMetric", "signature": "class LatentChangeMetric"}, {"doc": "Aggregated counters describing a single token's reasoning episode.", "kind": "class", "line": 648, "name": "ReasoningIterationStats", "signature": "class ReasoningIterationStats"}, {"doc": "Aggregated statistics over the full generation episode.", "kind": "class", "line": 661, "name": "GenerationReasoningSummary", "signature": "class GenerationReasoningSummary"}, {"doc": "Persists the high-level latent state across consecutive emitted tokens.\n\nThe cache is reset whenever its age in tokens reaches the configured\npersistence horizon, at which point the next reasoning episode begins\nwith a zero high-level state. This is the temporal-sparsity mechanism:\nexpensive full-stack passes are amortized across multiple emissions.", "kind": "class", "line": 684, "name": "SparseHighLevelStateCache", "signature": "class SparseHighLevelStateCache"}, {"doc": "Parameter-free hierarchical recursive reasoning over a trained stack.\n\nThe reasoner does not own any learnable parameters. It treats the trained\nTopoGPT2 transformer layers as a deterministic recurrent step function\nand composes them into a two-speed damped fixed-point iteration that\nmirrors HRM, while never violating the activation distribution the\ntrained layers expect.\n\nAlgorithm per emitted token:\n\n    1. Run the standard full forward pass once to obtain the baseline\n       residual-stream latent z_base and the per-layer kv caches that\n       will cross the token boundary. z_base is the trained model's\n       native answer for this position.\n    2. If recursion is disabled or both iteration budgets are zero,\n       return z_base unchanged.\n    3. Optionally warm-start z by adding a fraction of the cached\n       refinement vector from previous tokens (sparse temporal memory).\n    4. Hierarchical refinement, all in residual-stream space:\n          for h_step in range(max_high_level_iters):\n              for l_step in range(max_low_level_iters):\n                  z <- z + low_level_step * (W_low(z) - z)\n              z <- z + high_level_step * (W_high(z) - z)\n       where W_low and W_high are the last low_level_window and\n       high_level_window trained layers respectively, invoked with the\n       prefix kv cache treated as immutable. Each update is damped, so\n       layer inputs remain close to the trained residual-stream\n       distribution.\n    5. Hard divergence guard: if the iterated latent drifts farther\n       from the baseline than max_drift_relative, reset to the baseline\n       and abort thinking for this token. This eliminates the\n       catastrophic-attractor failure mode without retraining.\n    6. Attractor halting per loop, plus a global cap on total layer\n       invocations.\n\nThe cached refinement returned to the sparse cache is z_final - z_base,\na small residual-stream displacement that persists across configurable\nhorizons to amortize thinking effort over multiple tokens.", "kind": "class", "line": 727, "name": "HierarchicalRecursiveReasoner", "signature": "class HierarchicalRecursiveReasoner"}, {"doc": "Applies temperature, repetition penalty, top-k filtering and multinomial draw.", "kind": "class", "line": 935, "name": "LogitsSampler", "signature": "class LogitsSampler"}, {"doc": "Immutable sampling parameters consumed by the generation engine.", "kind": "class", "line": 963, "name": "SamplingPolicy", "signature": "class SamplingPolicy"}, {"doc": "Quantitative summary of a single generation call.", "kind": "class", "line": 985, "name": "GenerationReport", "signature": "class GenerationReport"}, {"doc": "Runs autoregressive sampling driven by hierarchical recursive reasoning.\n\nThe engine reimplements the prompt encoding and token emission loop so\nthat the per-token latent state can be intercepted before final norm and\nLM-head projection. The intercepted state is handed to a\nHierarchicalRecursiveReasoner, which iterates the trained layer stack in\na two-speed loop until the attractor is reached. The final stabilized\nlatent is then projected to logits and sampled in the standard fashion.", "kind": "class", "line": 1000, "name": "HRMGenerationEngine", "signature": "class HRMGenerationEngine"}, {"doc": "Prints a GenerationReport to stdout using settings-defined formatting.", "kind": "class", "line": 1189, "name": "ResultRenderer", "signature": "class ResultRenderer"}, {"doc": "Orchestrator wiring loader, builder, reasoner, engine and renderer.", "kind": "class", "line": 1230, "name": "HRMInferencePipeline", "signature": "class HRMInferencePipeline"}, {"doc": "Translates command-line arguments into an HRMInferenceSettings instance.", "kind": "class", "line": 1283, "name": "CliArgumentParser", "signature": "class CliArgumentParser"}, {"doc": "CLI entry point. Returns a process exit code.", "kind": "method", "line": 1494, "name": "main", "signature": "def main(argv)"}, {"doc": "Return the architecture preset table indexed by scale name.", "kind": "method", "line": 221, "name": "scale_presets", "signature": "def scale_presets()"}, {"doc": "Return the resolved preset for the configured model scale.", "kind": "method", "line": 234, "name": "preset", "signature": "def preset(self)"}, {"doc": "Raise ValueError if any setting falls outside its safety bounds.", "kind": "method", "line": 244, "name": "validate", "signature": "def validate(self)"}, {"doc": "Return a configured Logger with a single deduplicated stdout handler.", "kind": "method", "line": 347, "name": "build", "signature": "def build(settings)"}, {"doc": "Join parts under root and return the canonical resolved path.\n\nRaises ValueError if the resolved path escapes root.", "kind": "method", "line": 367, "name": "resolve_under", "signature": "def resolve_under(root)"}, {"doc": "Validate path points to an existing regular file with the expected suffix.", "kind": "method", "line": 383, "name": "require_existing_file", "signature": "def require_existing_file(path, expected_suffix)"}, {"kind": "method", "line": 400, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Return the topogpt3.train module which re-exports model symbols.", "kind": "method", "line": 404, "name": "load", "signature": "def load(self)"}, {"kind": "method", "line": 416, "name": "__init__", "signature": "def __init__(self, settings)"}, {"doc": "Directory holding the active checkpoint slot.", "kind": "method", "line": 424, "name": "slot_dir", "signature": "def slot_dir(self)"}, {"doc": "Resolved path to the safetensors weights file inside the slot.", "kind": "method", "line": 428, "name": "model_file", "signature": "def model_file(self)"}, {"doc": "Resolved path to the JSON training-state file inside the slot.", "kind": "method", "line": 434, "name": "state_file", "signature": "def state_file(self)"}, {"doc": "Verify weights exist and the on-disk size lies within safety bounds.", "kind": "method", "line": 440, "name": "assert_ready", "signature": "def assert_ready(self)"}, {"kind": "method", "line": 462, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Recover N_KV_HEADS used at training by inspecting the k_proj shape.\n\nReturns None when the probe key is absent, signalling the caller to\nfall back to scale defaults rather than guess.", "kind": "method", "line": 466, "name": "detect_n_kv_heads", "signature": "def detect_n_kv_heads(self, weights_path, d_model, n_heads)"}, {"kind": "method", "line": 505, "name": "__init__", "signature": "def __init__(self, settings, source_module, logger)"}, {"doc": "Return a TopoGPT2Config dataclass ready to instantiate the model.", "kind": "method", "line": 511, "name": "build", "signature": "def build(self, n_kv_heads, vocab_size)"}, {"kind": "method", "line": 536, "name": "__init__", "signature": "def __init__(self, settings, source_module)"}, {"doc": "Return an instance of BPETokenizer bound to the configured encoding.", "kind": "method", "line": 540, "name": "build", "signature": "def build(self)"}, {"kind": "method", "line": 549, "name": "__init__", "signature": "def __init__(self, settings, source_module, logger)"}, {"doc": "Patch QuaternionSpectralLayer to use the 3-multiply Gauss contract.", "kind": "method", "line": 555, "name": "apply_if_enabled", "signature": "def apply_if_enabled(self)"}, {"kind": "method", "line": 567, "name": "__init__", "signature": "def __init__(self, settings, source_module, logger)"}, {"doc": "Build the TopoGPT2 graph, load weights into it, and return it in eval mode.", "kind": "method", "line": 573, "name": "assemble", "signature": "def assemble(self, aligned_cfg, paths)"}, {"kind": "method", "line": 604, "name": "__init__", "signature": "def __init__(self, settings, source_module, logger)"}, {"doc": "Seed all relevant RNGs using the model package helper when available.", "kind": "method", "line": 610, "name": "apply", "signature": "def apply(self)"}, {"kind": "method", "line": 627, "name": "__init__", "signature": "def __init__(self, epsilon_floor)"}, {"doc": "Return ||current - previous|| / max(||previous||, epsilon_floor).", "kind": "method", "line": 632, "name": "relative_change", "signature": "def relative_change(self, current, previous)"}, {"doc": "Fold a per-token sample into the running totals.", "kind": "method", "line": 671, "name": "absorb", "signature": "def absorb(self, sample)"}, {"kind": "method", "line": 693, "name": "__init__", "signature": "def __init__(self, persist_tokens)"}, {"doc": "Return the cached high-level state or a zeroed one when stale.\n\nThe boolean flag indicates whether the returned tensor came from a\nlive cache hit (True) or a fresh zero initialization (False).", "kind": "method", "line": 700, "name": "get_or_init", "signature": "def get_or_init(self, reference)"}, {"doc": "Store a fresh high-level state and increment the cache age.", "kind": "method", "line": 716, "name": "commit", "signature": "def commit(self, new_state)"}, {"doc": "Drop any cached state and reset the age counter.", "kind": "method", "line": 721, "name": "invalidate", "signature": "def invalidate(self)"}, {"kind": "method", "line": 768, "name": "__init__", "signature": "def __init__(self, layers, final_norm, reasoning_config, logger)"}, {"doc": "Return the number of trained transformer layers.", "kind": "method", "line": 789, "name": "num_layers", "signature": "def num_layers(self)"}, {"doc": "Forward z_in through every layer using base_kvs as immutable prefix cache.\n\nReturns the layer-stack output and the freshly produced per-layer kv\ncaches that incorporate the K and V derived from z_in.", "kind": "method", "line": 793, "name": "_full_pass", "signature": "def _full_pass(self, z_in, base_kvs)"}, {"doc": "Forward z_in through the trailing `window` layers only.\n\nThe per-layer kv caches produced during this read-only pass are\ndiscarded; only the baseline pass's committed kvs cross the token\nboundary, preserving cache consistency across thinking iterations.", "kind": "method", "line": 808, "name": "_window_pass", "signature": "def _window_pass(self, z_in, base_kvs, window)"}, {"doc": "Run hierarchical recursive thinking for a single emission step.\n\nArgs:\n    z_initial: token embedding of the new position, shape [B, 1, D].\n    base_kvs: per-layer kv cache for all previously emitted tokens,\n        treated as immutable during thinking iterations.\n    cached_refinement: persistent refinement displacement from prior\n        tokens, or None to skip the warm start.\n\nReturns:\n    A tuple (z_final, committed_kvs, refinement_for_cache, stats):\n        z_final is the latent state about to enter the final norm\n        and lm head; committed_kvs is the new per-layer kv cache\n        including this token's K and V from the baseline pass;\n        refinement_for_cache is z_final - z_baseline, to be\n        persisted across tokens; stats holds the loop counters.", "kind": "method", "line": 827, "name": "reason", "signature": "def reason(self, z_initial, base_kvs, cached_refinement)"}, {"kind": "method", "line": 938, "name": "__init__", "signature": "def __init__(self, logger)"}, {"doc": "Return a sampled token id tensor of shape [B, 1] from raw logits [B, V].", "kind": "method", "line": 941, "name": "sample", "signature": "def sample(self, logits, token_history, temperature, top_k, repetition_penalty)"}, {"doc": "Construct a SamplingPolicy from inference settings.", "kind": "method", "line": 973, "name": "from_settings", "signature": "def from_settings(cls, settings)"}, {"doc": "Return throughput in tokens/sec, clamped to avoid divide-by-zero.", "kind": "method", "line": 995, "name": "tokens_per_second", "signature": "def tokens_per_second(self, elapsed_floor)"}, {"kind": "method", "line": 1011, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Run the prompt through the full stack once, returning the final\nhidden state of the last position, the per-layer base kv caches that\ncover all prompt tokens except the last one, and the embedding of the\nlast prompt token as the seed for the first reasoning episode.", "kind": "method", "line": 1016, "name": "_encode_prompt", "signature": "def _encode_prompt(self, model, prompt_ids)"}, {"doc": "Generate a completion for prompt and return a GenerationReport.", "kind": "method", "line": 1048, "name": "run", "signature": "def run(self, model, tokenizer, prompt, policy)"}, {"kind": "method", "line": 1192, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Emit a banner with prompt, completion, throughput and reasoning stats.", "kind": "method", "line": 1196, "name": "render", "signature": "def render(self, report)"}, {"kind": "method", "line": 1233, "name": "__init__", "signature": "def __init__(self, settings, logger)"}, {"doc": "Run the full inference pipeline end-to-end and return the report.", "kind": "method", "line": 1239, "name": "execute", "signature": "def execute(self)"}, {"doc": "Return the configured argparse.ArgumentParser.", "kind": "method", "line": 1287, "name": "build_parser", "signature": "def build_parser()"}, {"doc": "Parse argv (or sys.argv) and return a populated HRMInferenceSettings.", "kind": "method", "line": 1448, "name": "parse", "signature": "def parse(argv)"}]}, {"id": "topogpt3/jlens.py", "kind": "module", "label": "jlens.py", "language": "py", "sha256": "4aefc982cee01936", "symbol_count": 29, "symbols": [{"doc": "Centralized configuration for Jacobian lens fitting.\n\nEvery value consumed downstream resides here. Adding a new tunable means\nextending this class; no other module should embed literals.", "kind": "class", "line": 37, "name": "TopoGPT3JLensFitConfig", "signature": "class TopoGPT3JLensFitConfig"}, {"doc": "Centralized configuration for Jacobian lens application.\n\nEvery value consumed downstream resides here. Adding a new tunable means\nextending this class; no other module should embed literals.", "kind": "class", "line": 56, "name": "TopoGPT3JLensAppConfig", "signature": "class TopoGPT3JLensAppConfig"}, {"doc": "Captures residual-stream tensors at the given block indices.\n\nRegisters a forward hook on each requested block on ``__enter__`` and\nremoves them on ``__exit__``. On the next forward pass each block's output\nis stored in ``activations``, keyed by block index. Stored tensors are\nnot detached, so they can be passed straight to ``torch.autograd.grad``.\n\nArgs:\n    blocks: The sequence of residual blocks (e.g. ``model.layers``).\n    at: Block indices to record at.\n    start_graph_at: If given, the captured tensor at this index is marked\n        ``requires_grad_(True)`` before downstream blocks see it. When the\n        model's parameters all have ``requires_grad=False``, this makes the\n        captured residual the leaf that roots the autograd graph, so the\n        retained graph spans only this block onward.", "kind": "class", "line": 69, "name": "ActivationRecorder", "signature": "class ActivationRecorder"}, {"doc": "Boolean mask over sequence positions to include in the Jacobian average.\n\nEarly positions are dominated by attention-sink behaviour and the final\nposition has no next-token target, so both are excluded.\n\nArgs:\n    seq_len: Length of the tokenized prompt.\n    skip_first: Number of leading positions to exclude.\n\nReturns:\n    Boolean tensor of shape ``[seq_len]``.\n\nRaises:\n    ValueError: If ``skip_first`` is negative or the prompt is too short to\n        leave any valid positions.", "kind": "method", "line": 132, "name": "valid_position_mask", "signature": "def valid_position_mask(seq_len)"}, {"doc": "Resolve None/negative layer indices, bounds-check, enforce source < target.", "kind": "method", "line": 162, "name": "_check_layer_indices", "signature": "def _check_layer_indices(source_layers, target_layer, n_layers)"}, {"doc": "Compute the per-layer Jacobian estimator ``J_l`` for one prompt.\n\nRuns one forward pass on the prompt replicated ``dim_batch`` times along\nthe batch axis, retains the graph, then runs ``ceil(d_model / dim_batch)``\nbackward passes against it. Each backward computes ``dim_batch`` rows of\n``J_l`` at once: batch element ``b`` carries a one-hot cotangent at output\ndimension ``dim_start + b``, set at every valid target position.\n\nArgs:\n    model: The model to compute Jacobians for.\n    prompt: Input text.\n    source_layers: Layer indices ``l`` to compute ``J_l`` at.\n    target_layer: Layer to take gradients with respect to. Defaults to the\n        final layer; negative indices count from the end.\n    dim_batch: Output dimensions computed per backward pass.\n    max_seq_len: Truncate the prompt to this many tokens.\n    skip_first: Leading positions to exclude.\n\nReturns:\n    ``(jacobians, seq_len, n_valid_positions)``. ``jacobians`` maps each\n    source layer to a ``[d_model, d_model]`` fp32 CPU tensor.", "kind": "method", "line": 187, "name": "jacobian_for_prompt", "signature": "def jacobian_for_prompt(model, prompt, source_layers)"}, {"doc": "``torch.save`` to a temp file then ``os.replace`` so a crash never\nleaves a half-written checkpoint.", "kind": "method", "line": 283, "name": "_atomic_save", "signature": "def _atomic_save(obj, path)"}, {"doc": "Fit ``J_l`` over a list of prompts and return a JacobianLens.\n\nPer-prompt Jacobians from ``jacobian_for_prompt`` are accumulated as a\nrunning mean. If ``checkpoint_path`` is set, the running sum is written\nevery ``checkpoint_every`` prompts (atomic) and resumed from on restart.\n\nArgs:\n    model: The model to fit on.\n    prompts: Text prompts to average over.\n    source_layers: Layers to fit at. Defaults to every layer below\n        ``target_layer``; negative indices count from the end.\n    target_layer: See ``jacobian_for_prompt``.\n    dim_batch: See ``jacobian_for_prompt``.\n    max_seq_len: Truncate each prompt to this many tokens.\n    skip_first: See ``jacobian_for_prompt``.\n    checkpoint_path: If set, write a resumable checkpoint here.\n    checkpoint_every: Write checkpoint every N prompts (default 1).\n    resume: If True and checkpoint_path exists, resume from it.\n\nReturns:\n    The fitted JacobianLens.\n\nRaises:\n    ValueError: If no prompts are long enough to fit on, or if checkpoint\n        settings mismatch.", "kind": "method", "line": 291, "name": "fit", "signature": "def fit(model, prompts)"}, {"doc": "A fitted Jacobian lens: per-layer ``J_l`` matrices and the readout method.\n\nAttributes:\n    jacobians: ``{layer_index: Tensor[d_model, d_model]}``. Each ``J_l``\n        maps the residual at layer ``l`` into the final-layer basis.\n    source_layers: Sorted list of fitted layer indices.\n    n_prompts: Number of prompts the lens was averaged over.\n    d_model: Residual-stream width.", "kind": "class", "line": 459, "name": "JacobianLens", "signature": "class JacobianLens"}, {"doc": "Text-format slice data: top-K token predictions per (position, layer).\n\n``layers`` always includes the model's final layer (the actual model\noutput) so divergences from lens-transported earlier layers are visible.\n\nAttributes:\n    seq_len: Number of token positions in the slice.\n    layers: Layer indices shown (includes final layer).\n    prompt: The input prompt text.\n    input_ids: Tensor ``[1, seq_len]`` of token IDs.\n    token_strs: Decoded strings for each token position.\n    top_ids: ``[seq_len, n_layers, top_n]`` top token IDs per cell.\n    top_probs: ``[seq_len, n_layers, top_n]`` softmax probabilities.\n    top_token_strs: ``[seq_len, n_layers, top_n]`` decoded token strings\n        for each prediction. Empty string if tokenizer was unavailable.", "kind": "class", "line": 664, "name": "SliceData", "signature": "class SliceData"}, {"doc": "Compute a position x layer slice of top-K token predictions.\n\nFor each layer in the fitted lens, projects the residual at each position\nthrough the Jacobian into the final-layer basis, then unembeds to get\nlogits and softmax probabilities. Returns the top-N predicted token IDs\nand their probabilities per (position, layer) cell.\n\nArgs:\n    model: The model to read out from.\n    lens: A fitted JacobianLens.\n    prompt: Input text.\n    top_n: Top tokens to keep per (position, layer) cell.\n    max_seq_len: Truncate the prompt to this many tokens.\n\nReturns:\n    A SliceData instance with arrays indexed ``[seq_len, n_layers, top_n]``.", "kind": "method", "line": 705, "name": "compute_slice", "signature": "def compute_slice(model, lens, prompt)"}, {"doc": "Render a SliceData as a readable text table showing decoded words.\n\nFor each token position, shows what each layer predicts as the next token.\nThe first column shows the actual input token; subsequent columns show the\ntop-1 prediction at each layer with its softmax probability. Token strings\nare read from ``slice_data.top_token_strs`` (always populated by\n``compute_slice``).\n\nArgs:\n    slice_data: The slice to render.\n    tokenizer: Legacy parameter, ignored. Top token strings are already\n        stored in ``slice_data.top_token_strs``.\n    n_cols: Number of layer columns to show (default 3).\n\nReturns:\n    A multi-line string table.", "kind": "method", "line": 789, "name": "text_slice", "signature": "def text_slice(slice_data, tokenizer, n_cols)"}, {"doc": "Run a full jacobian lens demo loading real weights from checkpoint.", "kind": "method", "line": 842, "name": "_demo_jlens", "signature": "def _demo_jlens()"}, {"kind": "method", "line": 87, "name": "__init__", "signature": "def __init__(self, blocks, at)"}, {"kind": "method", "line": 102, "name": "_make_hook", "signature": "def _make_hook(self, index)"}, {"kind": "method", "line": 113, "name": "__enter__", "signature": "def __enter__(self)"}, {"kind": "method", "line": 126, "name": "__exit__", "signature": "def __exit__(self)"}, {"kind": "method", "line": 378, "name": "write_checkpoint", "signature": "def write_checkpoint()"}, {"kind": "method", "line": 470, "name": "__init__", "signature": "def __init__(self, jacobians)"}, {"kind": "method", "line": 482, "name": "__repr__", "signature": "def __repr__(self)"}, {"doc": "Save to ``path``. Jacobians are stored as ``dtype`` (default fp16).", "kind": "method", "line": 489, "name": "save", "signature": "def save(self, path)"}, {"doc": "Load a lens previously written by ``save``.", "kind": "method", "line": 504, "name": "load", "signature": "def load(cls, path)"}, {"doc": "Load a lens from a local file, a local directory, or a HuggingFace\nHub ``repo_id``.\n\n``filename`` is the path inside the directory or repo; ignored when\n``name_or_path`` is itself a file. ``revision`` selects a Hub branch,\ntag, or commit.", "kind": "method", "line": 519, "name": "from_pretrained", "signature": "def from_pretrained(cls, name_or_path)"}, {"doc": "Combine lenses fitted on disjoint prompt subsets into one\n(``n_prompts``-weighted mean of the inputs).\n\nArgs:\n    lenses: Lenses to merge. Must agree on ``source_layers`` and\n        ``d_model``.\n\nRaises:\n    ValueError: If ``lenses`` is empty or the inputs disagree on shape.", "kind": "method", "line": 543, "name": "merge", "signature": "def merge(cls, lenses)"}, {"doc": "Map a residual at ``layer`` into the final-layer basis: ``J_l @ h``.\n\nArgs:\n    residual: Tensor of shape ``[..., d_model]``.\n    layer: Source layer index (must be in ``source_layers``).", "kind": "method", "line": 574, "name": "transport", "signature": "def transport(self, residual, layer)"}, {"doc": "Run ``model`` on ``prompt`` and return lens logits at ``positions``.\n\nArgs:\n    model: The model to read out from.\n    prompt: Input text.\n    layers: Layers to read out at. Defaults to all of\n        ``source_layers``. Must be a subset of ``source_layers`` when\n        ``use_jacobian`` is True.\n    positions: Token positions to read out (Python indexing into the\n        sequence; negative indices count from the end). None returns\n        every position.\n    max_seq_len: Truncate the prompt to this many tokens.\n    use_jacobian: If False, skip the ``J_l`` transport (vanilla\n        logit-lens baseline).\n\nReturns:\n    A triple ``(lens_logits, model_logits, input_ids)``. ``lens_logits``\n    maps each requested layer to a ``[n_positions, vocab_size]`` tensor;\n    ``model_logits`` is the model's actual final-layer logits at the\n    same positions (same shape).\n\nRaises:\n    ValueError: If any requested layer is out of range for the model,\n        or (with use_jacobian) not in source_layers.", "kind": "method", "line": 585, "name": "apply", "signature": "def apply(self, model, prompt)"}, {"kind": "method", "line": 692, "name": "__post_init__", "signature": "def __post_init__(self)"}, {"kind": "method", "line": 105, "name": "hook", "signature": "def hook(module, inputs, output)"}, {"kind": "method", "line": 646, "name": "select", "signature": "def select(layer)"}]}, {"id": "topogpt3/lens_model.py", "kind": "module", "label": "lens_model.py", "language": "py", "sha256": "47fabe41f0bbfa6b", "symbol_count": 29, "symbols": [{"doc": "What the lens needs from a model.\n\nAttributes:\n    n_layers: Number of residual blocks.\n    d_model: Residual-stream width.\n    layers: The residual blocks, indexable by integer; what\n        ActivationRecorder hooks.\n    tokenizer: Tokenizer used by the visualisation helpers; must provide\n        ``decode(token_ids) -> str``. Fitting and apply() never touch it.", "kind": "class", "line": 23, "name": "LensModel", "signature": "class LensModel(Protocol)"}, {"doc": "Centralized configuration for the TopoGPT3 lens model adapter.\n\nEvery value consumed downstream resides here. Adding a new tunable means\nextending this class; no other module should embed literals.", "kind": "class", "line": 59, "name": "TopoGPT3LensConfig", "signature": "class TopoGPT3LensConfig"}, {"doc": "Runs the residual block stack only (no final norm, no LM head).\n\nThis is the forward subgraph that ActivationRecorder hooks capture.\nExtracted from TopoGPT2.forward() to expose the residual stream for\nJacobian lens fitting and application.", "kind": "class", "line": 142, "name": "_TopoGPT3ResidualForward", "signature": "class _TopoGPT3ResidualForward(Module)"}, {"doc": "LensModel adapter over a loaded TopoGPT2 model.\n\nWraps a TopoGPT2 instance and implements the LensModel protocol for use\nwith ActivationRecorder, JacobianLens fitting, and apply().\n\nThe adapter owns no parameters --- all weights live in the wrapped model.\nCall ``.eval()`` and set ``requires_grad_(False)`` on the wrapped model\nbefore fitting.", "kind": "class", "line": 161, "name": "TopoGPT3LensModel", "signature": "class TopoGPT3LensModel(Module)"}, {"doc": "A tiny CPU-only decoder for end-to-end tests.\n\nImplements the LensModel protocol indirectly (wrapped by\nTopoGPT3LensModel). Residual blocks are ``h + 0.1 * linear(h)``:\nthe small gain keeps the Jacobian well-conditioned so the late-layer\n``diag(J) ~= 1`` property holds.", "kind": "class", "line": 306, "name": "TinyDecoder", "signature": "class TinyDecoder(Module)"}, {"kind": "class", "line": 359, "name": "_ResidualBlock", "signature": "class _ResidualBlock(Module)"}, {"doc": "Tokenize ``text`` to ``input_ids`` of shape ``[1, seq_len]`` on the\nmodel's input device.", "kind": "method", "line": 40, "name": "encode", "signature": "def encode(self, text)"}, {"doc": "Run the residual stack on ``input_ids`` (no LM head). Must build an\nautograd graph through layers when grad is enabled, and must be\ndeterministic across batch elements (eval mode, dropout off) --- the\nfitting estimator replicates the prompt along the batch axis.", "kind": "method", "line": 45, "name": "forward", "signature": "def forward(self, input_ids)"}, {"doc": "Map a residual-stream tensor ``[..., d_model]`` to logits\n``[..., vocab_size]`` (final norm + LM head).", "kind": "method", "line": 52, "name": "unembed", "signature": "def unembed(self, residual)"}, {"doc": "Construct a lens config from a TopoGPT2Config dataclass.", "kind": "method", "line": 84, "name": "from_topogpt2_config", "signature": "def from_topogpt2_config(cls, cfg)"}, {"doc": "Probe a checkpoint directory and infer lens config from state.json.\n\nArgs:\n    checkpoint_dir: Path to the checkpoint slot directory.\n    state_filename: JSON file containing training config.\n\nReturns:\n    A TopoGPT3LensConfig matching the checkpoint.\n\nRaises:\n    FileNotFoundError: If state.json is missing.\n    ValueError: If required fields are absent from the state.", "kind": "method", "line": 104, "name": "probe_checkpoint", "signature": "def probe_checkpoint(cls, checkpoint_dir)"}, {"kind": "method", "line": 150, "name": "__init__", "signature": "def __init__(self, model)"}, {"kind": "method", "line": 154, "name": "forward", "signature": "def forward(self, input_ids)"}, {"kind": "method", "line": 172, "name": "__init__", "signature": "def __init__(self, model, tokenizer)"}, {"kind": "method", "line": 184, "name": "n_layers", "signature": "def n_layers(self)"}, {"kind": "method", "line": 188, "name": "d_model", "signature": "def d_model(self)"}, {"kind": "method", "line": 192, "name": "layers", "signature": "def layers(self)"}, {"kind": "method", "line": 196, "name": "tokenizer", "signature": "def tokenizer(self)"}, {"kind": "method", "line": 200, "name": "tokenizer", "signature": "def tokenizer(self, tok)"}, {"kind": "method", "line": 204, "name": "input_device", "signature": "def input_device(self)"}, {"kind": "method", "line": 210, "name": "input_device", "signature": "def input_device(self, device)"}, {"doc": "Tokenize text to input_ids of shape ``[1, seq_len]``.\n\nUses BPETokenizer if available, otherwise falls back to a byte-level\nencoding compatible with GPT-2 BPE tokenization.", "kind": "method", "line": 213, "name": "encode", "signature": "def encode(self, text)"}, {"doc": "Run the residual stack on ``input_ids``.\n\nReturns hidden states of shape ``[batch, seq_len, d_model]``\n(pre-final-norm, pre-LM-head). The autograd graph is retained through\nall layers when grad is enabled.", "kind": "method", "line": 228, "name": "forward", "signature": "def forward(self, input_ids)"}, {"doc": "Map residual ``[..., d_model]`` to logits ``[..., vocab_size]``.\n\nApplies the model's final norm and LM head projection.", "kind": "method", "line": 237, "name": "unembed", "signature": "def unembed(self, residual)"}, {"doc": "Build a TopoGPT3LensModel from a checkpoint directory.\n\nProbes state.json for configuration, instantiates the model, loads\nsafetensors weights, and wraps the result.\n\nArgs:\n    checkpoint_dir: Path to the checkpoint slot directory.\n    device: Target device. Defaults to cuda if available else cpu.\n    encoding: Tokenizer encoding name (passed to BPETokenizer).\n    strict: Whether to enforce strict state dict loading.\n\nReturns:\n    A TopoGPT3LensModel in eval mode with requires_grad_(False).\n\nRaises:\n    FileNotFoundError: If model.safetensors or state.json is missing.", "kind": "method", "line": 246, "name": "from_checkpoint", "signature": "def from_checkpoint(cls, checkpoint_dir)"}, {"kind": "method", "line": 315, "name": "__init__", "signature": "def __init__(self, n_layers, d_model, vocab_size, seed)"}, {"kind": "method", "line": 344, "name": "forward", "signature": "def forward(self, token_ids, past_kvs)"}, {"kind": "method", "line": 360, "name": "__init__", "signature": "def __init__(self, d_model)"}, {"kind": "method", "line": 366, "name": "forward", "signature": "def forward(self, x, past_kv)"}]}, {"doc": "TopoMerged: 6-tier curriculum merging TopoGPT3 code tasks with ExploitGym exploit tasks.  Tier structure: Phase 1 — Code foundation (from TopoGPT3): Tier 1: codealpaca            (128 seq, 2 epochs) — short instructions Tier 2: code_feedback         (192 seq, 2 epochs) — step-by-step explanations Tier 3: magicoder_evol        (256 seq, 3 epochs) — complex problems  Phase 2 — Exploit specialization (from ExploitGym): Tier 4: vulnerability_analysis (256 seq, 4 epochs) — analyze CVEs Tier 5: patch_analysis         (512 seq, 3 epochs) — understand patches Tier 6: exploit_development    (768 seq, 4 epochs) — write exploits  The model learns to code first, then applies coding skill to security.  Author: Gris Iscomeback License: GPL v3", "id": "topogpt3/merged_config.py", "kind": "module", "label": "merged_config.py", "language": "py", "sha256": "6add51e0e992f8db", "symbol_count": 6, "symbols": [{"doc": "Configuration for TopoMerged: 6-tier code+exploit curriculum.", "kind": "class", "line": 32, "name": "TopoMergedConfig", "signature": "class TopoMergedConfig"}, {"kind": "method", "line": 139, "name": "build_topogpt2_config", "signature": "def build_topogpt2_config(self, max_seq_len, attn_window)"}, {"kind": "method", "line": 161, "name": "build_exploit_loader_config", "signature": "def build_exploit_loader_config(self)"}, {"doc": "Returns True if the tier uses TopoGPT3 (HuggingFace) data.", "kind": "method", "line": 175, "name": "is_code_tier", "signature": "def is_code_tier(self, tier_index)"}, {"doc": "Returns True if the tier uses ExploitGym data.", "kind": "method", "line": 179, "name": "is_exploit_tier", "signature": "def is_exploit_tier(self, tier_index)"}, {"doc": "Convert global tier index to exploitgym-local index (0-2).", "kind": "method", "line": 183, "name": "exploit_tier_index", "signature": "def exploit_tier_index(self, tier_index)"}]}, {"doc": "TopoGPT2: Quaternion-Enhanced Topological Transformer Language Model  Author: Gris Iscomeback Email: grisiscomeback@gmail.com License: GPL v3  Mejoras sobre topogpt.py: - Álgebra de cuaterniones completa (QuaternionLinear, QuaternionSpectralLayer) con producto de Hamilton en el dominio de frecuencia para capturar la espectrografía de los datos con kernels reales e imaginarios cruzados. - SpectralAutoencoder: encoder/decoder espectral que comprime y reconstruye las representaciones en el dominio de frecuencia. - QuaternionTorusBrain VECTORIZADA (sin bucles sobre seq_len): proyección geométrica sobre el toro con asignación blanda usando distancias circulares, message-passing con rotaciones de cuaterniones. - 8 nodos (RADIAL=2 × ANGULAR=4), 4 ángulos, 2 radiales (spec del usuario). - Rotary Position Embeddings (RoPE). - Flash-attention (scaled_dot_product_attention de PyTorch 2.0+). - RMSNorm en lugar de LayerNorm (estilo LLaMA). - Tokenizador BPE via tiktoken (vocab GPT-2, 50k tokens). - Descargador de corpus: TinyStories, WikiText-103, raw file. - Entrenamiento con AMP (mixed precision) + acumulación de gradientes. - Presets de escala: micro, small, medium, gpt2.", "id": "topogpt3/model.py", "kind": "module", "label": "model.py", "language": "py", "sha256": "778e4b52ef6fea82", "symbol_count": 187, "symbols": [{"doc": "Configuración completa para TopoGPT2.", "kind": "class", "line": 56, "name": "TopoGPT2Config", "signature": "class TopoGPT2Config"}, {"kind": "method", "line": 194, "name": "setup_logger", "signature": "def setup_logger(name, level)"}, {"kind": "method", "line": 204, "name": "set_seed", "signature": "def set_seed(seed, device)"}, {"doc": "Operaciones de cuaterniones puras en PyTorch.\nRepresentación: [..., 4]  donde last dim = [w, x, y, z]\nq = w + x*i + y*j + z*k", "kind": "class", "line": 216, "name": "QuaternionOps", "signature": "class QuaternionOps"}, {"doc": "Capa lineal con pesos cuaterniones.\n\nImplementa la multiplicación W * x en el álgebra de cuaterniones:\n- W = Ww + Wx*i + Wy*j + Wz*k  (cuaternión de pesos)\n- x = xw + xx*i + xy*j + xz*k  (cuaternión de entrada)\n- out = W * x  (producto de Hamilton extendido a vectores)\n\nParámetros: 4 matrices reales de forma [out_q, in_q]", "kind": "class", "line": 255, "name": "QuaternionLinear", "signature": "class QuaternionLinear(Module)"}, {"doc": "Convolución espectral 2D con cuaterniones y producto de Hamilton completo.\n\nOperación en dominio de frecuencia:\n    P(k) = W(k) ⊗ X(k)  (producto de Hamilton de cuaterniones complejos)\n\nDonde:\n    X(k) = FFT2(x) con 4 canales cuaterniones [Xw, Xx, Xy, Xz]\n    W(k) = kernel complejo aprendible con componentes [Ww, Wx, Wy, Wz]\n\nReglas del producto de Hamilton en dominio de frecuencia:\n    Pw = Ww·Xw - Wx·Xx - Wy·Xy - Wz·Xz\n    Px = Ww·Xx + Wx·Xw + Wy·Xz - Wz·Xy\n    Py = Ww·Xy - Wx·Xz + Wy·Xw + Wz·Xx\n    Pz = Ww·Xz + Wx·Xy - Wy·Xx + Wz·Xw\n\nCada Wc es un kernel complejo (partes real e imaginaria independientes).", "kind": "class", "line": 300, "name": "QuaternionSpectralLayer", "signature": "class QuaternionSpectralLayer(Module)"}, {"doc": "Autoencoder espectral con cuaterniones.\n\nOpera en dos niveles:\n1. Espectral 1D sobre el vector de features (FFT sobre dim D_MODEL):\n   captura la espectrografía global del embedding.\n2. Espectral 2D sobre el grid del toro (QuaternionSpectralLayer):\n   captura correlaciones espaciales en la topología.\n\nDevuelve (latent, recon_loss) para regularización.", "kind": "class", "line": 387, "name": "SpectralAutoencoder", "signature": "class SpectralAutoencoder(Module)"}, {"doc": "Reemplaza el MLP en cada capa del transformer.\n\nPipeline (completamente vectorizado sobre batch Y secuencia):\n\n1. Flatten: [B, S, D] → [B·S, D]\n2. SpectralAutoencoder: filtrado espectral 1D + compresión cuaternión\n3. Proyección al toro:\n   - Calcula 2 ángulos (phi1, phi2) ∈ [-π, π]²\n   - Asignación blanda a los 8 nodos via distancia circular en el toro\n4. Construye grid de nodos: [B·S, N_NODES=8, D_MODEL]\n5. QuaternionSpectralLayer 2D sobre el grid [B·S, 4*D_QUAT, RADIAL, ANGULAR]\n6. Message-passing con rotaciones cuaterniones sobre el grafo toro\n7. Readout: atención sobre los 8 nodos → [B·S, D_MODEL]\n8. Reshape: [B·S, D] → [B, S, D]", "kind": "class", "line": 470, "name": "QuaternionTorusBrain", "signature": "class QuaternionTorusBrain(Module)"}, {"doc": "Rotary Position Embeddings (RoPE) - Su et al., 2021.\nCodifica la posicion como rotaciones del espacio de atencion,\nnaturalmente relativas y sin parametros extra.\n\nLas caches _cos/_sin se registran como buffers no-persistentes con\nnombres que no colisionan con checkpoints antiguos (que usaban\n'cos_cache'/'sin_cache'). Esto permite cambiar MAX_SEQ_LEN sin\nerrores de shape al cargar checkpoints previos.", "kind": "class", "line": 687, "name": "RotaryEmbedding", "signature": "class RotaryEmbedding(Module)"}, {"doc": "Root Mean Square Layer Normalization (sin bias). Más estable que LayerNorm.", "kind": "class", "line": 741, "name": "RMSNorm", "signature": "class RMSNorm(Module)"}, {"doc": "SwiGLU: SiLU(gate(x)) * up(x) -> down\nUsado en LLaMA 2/3, Qwen, Mistral en lugar de GELU-FFN.\nDimension interna: 8/3 * d_model (convención LLaMA, redondeada a múltiplo de 4).", "kind": "class", "line": 758, "name": "SwiGLU", "signature": "class SwiGLU(Module)"}, {"doc": "Mixture of Experts sobre la capa topologica.\n\nArquitectura (inspirada en DeepSeek-MoE / Mixtral):\n  - 1 experto compartido: QuaternionTorusBrain (siempre activo)\n  - N_EXPERTS expertos SwiGLU ligeros (activacion esparsa: Top-K por token)\n  - Router: Linear(D, N_EXPERTS) + softmax → top-K\n\nLoad-balancing loss (auxiliar): penaliza si un experto acapara todos los tokens.\nActiva MOE_TOP_K de N_EXPERTS expertos por token.\n\nSin MoE (MOE_ENABLED=False): se comporta como QuaternionTorusBrain puro.", "kind": "class", "line": 787, "name": "TopoMoEBrain", "signature": "class TopoMoEBrain(Module)"}, {"doc": "Multi-head attention con:\n- Flash Attention (scaled_dot_product_attention de PyTorch 2.0+)\n- Rotary Position Embeddings (RoPE)\n- GQA (Grouped Query Attention): N_KV_HEADS < N_HEADS, reduce VRAM de K/V\n- KV Cache para inferencia autoregresiva eficiente\n- Temperatura termodinámica aprendible", "kind": "class", "line": 892, "name": "MultiHeadAttention", "signature": "class MultiHeadAttention(Module)"}, {"doc": "Capa del transformer con TopoMoEBrain (TopoBrain + MoE SwiGLU experts).\n\nEsquema pre-norm (estilo LLaMA):\n    x = x + Attention_GQA(RMSNorm(x))\n    x = x + TopoMoEBrain(RMSNorm(x))", "kind": "class", "line": 996, "name": "TopoGPT2Layer", "signature": "class TopoGPT2Layer(Module)"}, {"doc": "TopoGPT2: Transformer de lenguaje con TopoBrain cuaternión-espectral.\n\nArquitectura:\n    Embedding de tokens + RoPE (en Attention)\n    N_LAYERS × TopoGPT2Layer (Attention + QuaternionTorusBrain)\n    RMSNorm final\n    Proyección a vocabulario (weight-tied con embeddings)", "kind": "class", "line": 1043, "name": "TopoGPT2", "signature": "class TopoGPT2(Module)"}, {"doc": "Wrapper alrededor de tiktoken (GPT-2 compatible).", "kind": "class", "line": 1260, "name": "BPETokenizer", "signature": "class BPETokenizer"}, {"doc": "Disk-cached manifest of text files found in a directory tree.", "kind": "class", "line": 1386, "name": "FileManifest", "signature": "class FileManifest"}, {"doc": "Tokenizes file paths into a memory-mapped numpy array on disk.\n\nUses incremental file reading and batched writing to avoid loading\nall tokens into RAM. Tokens are stored as raw int64 on disk and\naccessed via numpy memmap (OS-level paging, near-zero RAM footprint).", "kind": "class", "line": 1453, "name": "MemmapTokenizer", "signature": "class MemmapTokenizer"}, {"doc": "Memory-mapped token dataset for sequence-to-sequence LM training.\n\nThe token array is backed by a numpy memmap file on disk.\nOnly accessed slices are paged into RAM by the OS. The .copy()\nin __getitem__ ensures the returned torch.Tensor owns its memory,\nwhich is required for DataLoader collation with worker processes.", "kind": "class", "line": 1542, "name": "MappedTokenDataset", "signature": "class MappedTokenDataset(Dataset)"}, {"doc": "Filters low-quality files from the corpus based on multiple heuristics.", "kind": "class", "line": 1576, "name": "TextFilter", "signature": "class TextFilter"}, {"doc": "Tiered dataset that exposes short/medium/all files based on line count.\n\nWorks as a wrapper around MappedTokenDataset. Provides __getitem__ that\nonly samples from the active tier, avoiding dataset duplication.", "kind": "class", "line": 1680, "name": "CurriculumDataset", "signature": "class CurriculumDataset(Dataset)"}, {"doc": "Classify file paths into complexity tiers by line count.\n\nReturns dict: tier -> list of file indices in that tier.\nTier 0 = short (<=short lines), tier 1 = medium, tier 2 = all.", "kind": "method", "line": 1718, "name": "build_file_tiers", "signature": "def build_file_tiers(paths, short, med)"}, {"doc": "Trainer that dynamically adjusts MAX_SEQ_LEN across training phases.\n\nPhase schedule (configurable):\n    phase 0: seq_len=128, epochs=3\n    phase 1: seq_len=256, epochs=3\n    phase 2: seq_len=512, epochs=4\n\nEach phase rebuilds the DataLoader with the new sequence length.", "kind": "class", "line": 1744, "name": "ProgressiveSeqLenTrainer", "signature": "class ProgressiveSeqLenTrainer"}, {"doc": "Speculative decoding with a small draft model.\n\nDraft model uses SPEC_DECODE_DRAFT_SCALE (e.g. 'micro').\nThe draft generates K tokens, then the target model verifies them\nin a single forward pass. Accepted tokens are kept; rejected ones\ntrigger a fallback to the target model sampling.", "kind": "class", "line": 1820, "name": "SpeculativeDecoder", "signature": "class SpeculativeDecoder"}, {"doc": "Wrapper around nn.Embedding that applies dynamic quantization.\n\nApplies int8 quantization to the embedding weight matrix after loading.\nSupports both embed (int8) and FFN (int4) quantization modes.", "kind": "class", "line": 1935, "name": "QuantizedEmbedding", "signature": "class QuantizedEmbedding(Module)"}, {"doc": "Quantize embedding and lm_head layers for reduced VRAM usage.", "kind": "method", "line": 1975, "name": "apply_quantization", "signature": "def apply_quantization(model, config)"}, {"doc": "Extends TopoGPT2Trainer with curriculum + progressive seq len support.\n\nProvides:\n- Tokens cache for progressive sequence length rebuilding\n- Curriculum dataset wrapping (short / medium / all tiers)", "kind": "class", "line": 2003, "name": "CurriculumTrainer", "signature": "class CurriculumTrainer"}, {"doc": "Tokenize a single text string and write tokens to disk as raw int64.", "kind": "method", "line": 2143, "name": "_tokenize_text_to_memmap", "signature": "def _tokenize_text_to_memmap(text, tokenizer, path, max_tokens)"}, {"doc": "Gestiona checkpoints de forma acumulativa y segura.\n\nEstructura en disco:\n    checkpoints_topogpt2/\n      latest/\n        model.safetensors   <- pesos del modelo (formato seguro, sin pickle)\n        optimizer.pt        <- estado del optimizador (requiere .pt)\n        state.json          <- metadatos: epoch, step, historial, config\n      best/\n        model.safetensors\n        state.json\n      step_NNNNN/           <- snapshots periodicos (rotados)\n        model.safetensors\n        optimizer.pt\n        state.json\n\nEl historial se ACUMULA entre sesiones de entrenamiento: cada --resume\nagrega nuevas entradas a train_loss[], val_loss[], etc.", "kind": "class", "line": 2155, "name": "CheckpointManager", "signature": "class CheckpointManager"}, {"doc": "Entrenador acumulativo y resumible.\n\nCaracteristicas:\n- Checkpoint automatico en safetensors cada N minutos + cada epoch\n- Historial acumulativo entre sesiones (--resume agrega al historial existente)\n- Guarda el mejor modelo en checkpoints/best/ automaticamente\n- LR schedule: cosine con warmup relativo a los steps de ESTA sesion\n- Mixed Precision (AMP) + acumulacion de gradientes", "kind": "class", "line": 2388, "name": "TopoGPT2Trainer", "signature": "class TopoGPT2Trainer"}, {"doc": "Calcula todas las metricas del diagrama de fases de Book.md.\n\nTodas las metricas se derivan de cantidades medibles (pesos, gradientes):\n\ndelta  (δ): margen de discretizacion.  max|w - round(w)|\n            δ≈0 -> cristal;  δ≈0.49 -> vidrio frio\nkappa  (κ): numero de condicion de la covarianza del gradiente.\n            κ≈1 -> cristalino;  κ>>1 -> amorfo\nT_eff:      temperatura efectiva = (lr/2) * Var(gradiente).\n            T_eff→0 -> congelado; T_eff alto -> ruidoso\nalpha  (α): indice de pureza = -log(δ + ε).\n            α=20 -> perfecto; α<1 -> vidrio\nberry:      fase de Berry de los kernels espectrales imaginarios.\n            |berry|>π/2 con winding≠0 -> insulador topologico\nlc:         complejidad local = 1 - similitud coseno promedio entre filas.\nsp:         superposicion = correlacion promedio inter-fila de pesos.", "kind": "class", "line": 2681, "name": "MechanisticMetrics", "signature": "class MechanisticMetrics"}, {"doc": "Encuentra el ratio imaginario/real optimo para los kernels espectrales.\n\nAnalogia con main.py: evalua la transicion GOE→GUE en el espacio\nde kernels. Un ratio optimo promueve estructura topologica (insulador)\nvs estructura amorfa (vidrio).\n\nMetodo: calibra con un mini-batch y mide la varianza del gradiente\nen funcion del ratio. Ratios que minimizan la varianza de gradiente\n(maxima coherencia espectral) son preferibles.\n\nNo entrena: solo inicializa los kernels con distintos ratios y mide.\nTiempo tipico: < 30 segundos.", "kind": "class", "line": 2918, "name": "Phase0_KernelOptimizer", "signature": "class Phase0_KernelOptimizer"}, {"doc": "Encuentra el batch size optimo testando candidatos con pocos pasos.\n\nDe main.py: el batch size regula la temperatura del horno de cristalizacion.\nBatch sizes demasiado chicos -> ruido excesivo (vidrio frio).\nBatch sizes demasiado grandes -> sin presion annealing (amorfos).\nLa ventana optima empirica de main.py: [24, 128] para Strassen.\n\nPara LM, testeamos candidatos midiendo:\n- delta (δ): velocidad de descenso en prospect_steps pasos\n- T_eff: temperatura efectiva del gradiente\n\nTiempo tipico: < 2 minutos para 3 candidatos × 30 pasos.", "kind": "class", "line": 2993, "name": "Phase1_BatchProspector", "signature": "class Phase1_BatchProspector"}, {"doc": "Encuentra semillas prometedoras midiendo la trayectoria de delta.\n\nDe main.py: una semilla \"buena\" muestra delta descendente en los\nprimeros N pasos (enfriamiento). Una semilla \"mala\" se estanca en\nel plateau vidrioso (~0.49).\n\nCriterio de seleccion:\n1. Semillas con delta_velocity < 0 (enfriando) AND kappa bajo.\n2. Si no hay, semillas solo enfriando.\n3. Fallback: semilla con menor delta final.\n\nTiempo tipico: < 3 minutos para 5 semillas × 50 pasos.", "kind": "class", "line": 3076, "name": "Phase2_SeedMiner", "signature": "class Phase2_SeedMiner"}, {"doc": "Refinamiento post-entrenamiento mediante recocido simulado.\n\nDe main.py: despues de que el modelo converge, una fase de annealing\ncon criterio de aceptacion de Metropolis puede empujar los pesos\nhacia estados de menor energia libre (menor delta o mejor val_loss).\n\nAceptacion de Metropolis:\n    si Δloss < 0: siempre acepta (mejora)\n    si Δloss >= 0: acepta con prob exp(-Δloss / T)\n\nLa temperatura T decae exponencialmente: T(t) = T0 * cooling_rate^t\n\nAl rechazar: restaura el mejor estado conocido.\nSi se estanca: perturbacion termica (ruido gaussiano en pesos).\n\nTiempo: proporcional a refine_epochs (user-controlled).", "kind": "class", "line": 3158, "name": "Phase4_AnnealingRefiner", "signature": "class Phase4_AnnealingRefiner"}, {"doc": "Pipeline with curriculum learning and progressive sequence length.\n\nReplaces TopoPhasePipeline when --curriculum or --progressive-seq-len is set.\nHandles:\n- Text quality filtering before tokenization (via TextFilter)\n- Curriculum tiers (short/medium/all files)\n- Progressive MAX_SEQ_LEN across phases: 128->256->512\n- Tokens cached in memory for fast DataLoader rebuilding per phase", "kind": "class", "line": 3319, "name": "TopoPhasePipelineV2", "signature": "class TopoPhasePipelineV2"}, {"doc": "Orquesta las 5 fases de entrenamiento segun main.py + Book.md.\n\nFases:\n  0  Kernel ratio optimization  (GOE-GUE spectral calibration)\n  1  Batch size prospecting      (temperatura del horno de cristalizacion)\n  2  Seed mining                 (seleccion de semilla enfriante)\n  3  Full training               (entrenamiento principal con metricas)\n  4  Annealing refinement        (recocido simulado post-entrenamiento)\n\nLas fases 0-2 son rapidas (prospecting). La fase 3 es el grueso.\nLa fase 4 es opcional (--refine).\n\nPara no ser prohibitivo:\n  --prospect         activa fases 0, 1, 2 antes del entrenamiento\n  --refine-epochs N  activa fase 4 con N epocas de annealing\n  Sin flags: solo fase 3 (comportamiento original, identico a antes)", "kind": "class", "line": 3448, "name": "TopoPhasePipeline", "signature": "class TopoPhasePipeline"}, {"kind": "method", "line": 3570, "name": "main", "signature": "def main()"}, {"kind": "method", "line": 161, "name": "__post_init__", "signature": "def __post_init__(self)"}, {"doc": "Producto de Hamilton q1 ⊗ q2. Ambos [..., 4].", "kind": "method", "line": 224, "name": "hamilton_product", "signature": "def hamilton_product(q1, q2)"}, {"kind": "method", "line": 236, "name": "normalize", "signature": "def normalize(q, eps)"}, {"kind": "method", "line": 240, "name": "conjugate", "signature": "def conjugate(q)"}, {"doc": "Rota vector 3D v por cuaternión unitario q. v:[...,3] q:[...,4]", "kind": "method", "line": 245, "name": "rotate_vector", "signature": "def rotate_vector(v, q)"}, {"kind": "method", "line": 267, "name": "__init__", "signature": "def __init__(self, in_features, out_features, bias)"}, {"doc": "x: [..., in_features] → [..., out_features]", "kind": "method", "line": 283, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 320, "name": "__init__", "signature": "def __init__(self, in_q, out_q, grid_h, grid_w, init_scale)"}, {"kind": "method", "line": 339, "name": "_kernel", "signature": "def _kernel(self, c)"}, {"doc": "Suma sobre canales in_q: Y[b,o,h,w] = Σ_i W[i,o,h,w]·X[b,i,h,w]", "kind": "method", "line": 342, "name": "_contract", "signature": "def _contract(self, W, X)"}, {"doc": "x: [B, 4*in_q, H, W]  (4 canales cuaterniones sobre grid espacial)\n→ [B, 4*out_q, H, W]", "kind": "method", "line": 346, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 400, "name": "__init__", "signature": "def __init__(self, config)"}, {"doc": "Filtro espectral 1D: x[..., D] → filtrado[..., D]", "kind": "method", "line": 432, "name": "_filter1d", "signature": "def _filter1d(self, x, kr, ki)"}, {"doc": "x: [..., D_MODEL] → latent: [..., D_LAT]", "kind": "method", "line": 438, "name": "encode", "signature": "def encode(self, x)"}, {"doc": "z: [..., D_LAT] → recon: [..., D_MODEL]", "kind": "method", "line": 443, "name": "decode", "signature": "def decode(self, z)"}, {"doc": "Devuelve (latent, recon_loss)", "kind": "method", "line": 448, "name": "forward", "signature": "def forward(self, x)"}, {"doc": "Procesa el grid del toro con QuaternionSpectralLayer.\ngrid: [B, 4*D_QUAT, RADIAL, ANGULAR]  →  [B, 4*D_QUAT, RADIAL, ANGULAR]", "kind": "method", "line": 455, "name": "process_torus_grid", "signature": "def process_torus_grid(self, grid)"}, {"kind": "method", "line": 488, "name": "__init__", "signature": "def __init__(self, d_model, config)"}, {"doc": "Construye las aristas del grafo toro 2×4.\n\nNodos indexados como: node = r * N_ANGULAR + a\n  r ∈ [0, RADIAL-1], a ∈ [0, ANGULAR-1]\n\nAristas angulares: nodo ↔ nodo a la izquierda/derecha (periódico)\nAristas radiales:  nodo ↔ nodo del anillo interior/exterior", "kind": "method", "line": 528, "name": "_build_torus_graph", "signature": "def _build_torus_graph(self)"}, {"doc": "Asignación blanda de tokens a los 8 nodos del toro via distancia circular.\n\nphi1: [BS] ángulo angular ∈ [-π, π]\nphi2: [BS] ángulo radial ∈ [-π, π]\n→ weights: [BS, N_NODES]  (suma a 1, softmax de distancias negativas)", "kind": "method", "line": 562, "name": "_torus_soft_assign", "signature": "def _torus_soft_assign(self, phi1, phi2)"}, {"doc": "Message-passing VECTORIZADO con rotaciones cuaterniones.\nSin bucles Python: todas las aristas se procesan en paralelo.\n\nnode_feat: [BS, N_NODES, D_MODEL]\n→ [BS, N_NODES, D_MODEL]", "kind": "method", "line": 589, "name": "_message_passing", "signature": "def _message_passing(self, node_feat)"}, {"doc": "x: [B, S, D_MODEL]\n→ output: [B, S, D_MODEL], recon_loss: scalar", "kind": "method", "line": 626, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 699, "name": "__init__", "signature": "def __init__(self, d_head, max_seq_len, base)"}, {"kind": "method", "line": 705, "name": "_build_cache", "signature": "def _build_cache(self, seq_len)"}, {"kind": "method", "line": 713, "name": "_rotate_half", "signature": "def _rotate_half(self, x)"}, {"doc": "q, k: [B, n_heads, S_q/S_k, d_head]\noffset: posicion inicial (para KV cache: longitud del cache existente)\nAplica posiciones [offset .. offset+S-1] a q y k.", "kind": "method", "line": 717, "name": "forward", "signature": "def forward(self, q, k, seq_len, offset)"}, {"kind": "method", "line": 744, "name": "__init__", "signature": "def __init__(self, d_model, eps)"}, {"kind": "method", "line": 749, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 765, "name": "__init__", "signature": "def __init__(self, d_model, expansion, dropout)"}, {"kind": "method", "line": 779, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 802, "name": "__init__", "signature": "def __init__(self, d_model, config)"}, {"doc": "x: [N, D] donde N = B*S (tokens aplanados)\nRetorna:\nexpert_out: [N, D]  suma ponderada de top-K expertos\naux_loss:   escalar  load-balancing loss\nRouting vectorizado sin boolean indexing ni sincronizacion CUDA.\nUsa dispatch por indices agrupados (estilo Mixtral/DeepSeek) para\ncompatibilidad total con torch.utils.checkpoint.", "kind": "method", "line": 823, "name": "_route", "signature": "def _route(self, x)"}, {"doc": "x: [B, S, D]\n→ output: [B, S, D], aux_loss: escalar", "kind": "method", "line": 865, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 902, "name": "__init__", "signature": "def __init__(self, d_model, n_heads, config)"}, {"doc": "Args:\n    x:        [B, S, D]\n    is_causal: usar mascara causal\n    past_kv:  (K_cache, V_cache) de pasos anteriores o None\nReturns:\n    out:      [B, S, D]\n    kv_cache: (K, V) completos para cachear en generate()", "kind": "method", "line": 921, "name": "forward", "signature": "def forward(self, x, is_causal, past_kv)"}, {"kind": "method", "line": 1005, "name": "__init__", "signature": "def __init__(self, d_model, n_heads, config)"}, {"kind": "method", "line": 1014, "name": "_forward_impl", "signature": "def _forward_impl(self, x, past_kv)"}, {"doc": "Retorna (x_out, aux_loss, kv_cache).\nCon gradient checkpointing en training (solo cuando no hay KV cache).", "kind": "method", "line": 1023, "name": "forward", "signature": "def forward(self, x, past_kv)"}, {"kind": "method", "line": 1054, "name": "__init__", "signature": "def __init__(self, config)"}, {"kind": "method", "line": 1079, "name": "_init_weights", "signature": "def _init_weights(self)"}, {"doc": "token_ids: [B, S]  (enteros)\npast_kvs:  lista de (K, V) por capa, o None para entrenamiento\n→ logits: [B, S, VOCAB_SIZE], aux_loss: scalar, new_kvs: list[(K,V)]", "kind": "method", "line": 1086, "name": "forward", "signature": "def forward(self, token_ids, past_kvs)"}, {"doc": "Process long sequences with latent memory-token context compression.\n\nSplits `token_ids` [B, S] into segments of size MEMORY_SEGMENT_LEN.\nEach segment is processed with N_MEMORY_TOKENS prepended. The output\nat memory-token positions after segment k becomes the memory-state\ninput for segment k+1, compressing all prior context into a fixed-size\nlatent vector.\n\nReturns (logits [B, S, VOCAB_SIZE], aux_loss).", "kind": "method", "line": 1109, "name": "forward_with_memory", "signature": "def forward_with_memory(self, token_ids)"}, {"kind": "method", "line": 1159, "name": "count_params", "signature": "def count_params(self)"}, {"doc": "Autoregressive generation with KV cache and top-k sampling.\n\nArgs:\n    token_ids: [B, S_prompt] prompt tokens.\n    max_new_tokens: Maximum tokens to generate.\n    temperature: Sampling temperature (lower = more deterministic).\n    top_k: Top-k filtering (0 = disabled).\n    repetition_penalty: Penalty for repeating tokens (>1 = penalize).\n\nReturns:\n    [B, S_prompt + generated] full token sequence.", "kind": "method", "line": 1165, "name": "generate", "signature": "def generate(self, token_ids, max_new_tokens, temperature, top_k, repetition_penalty)"}, {"kind": "method", "line": 1216, "name": "generate_with_continuation", "signature": "def generate_with_continuation(self, token_ids, tokenizer, max_new_tokens, temperature, top_k, repetition_penalty, max_continuations, tail_lines)"}, {"kind": "method", "line": 1263, "name": "__init__", "signature": "def __init__(self, encoding)"}, {"kind": "method", "line": 1271, "name": "encode", "signature": "def encode(self, text)"}, {"kind": "method", "line": 1274, "name": "decode", "signature": "def decode(self, tokens)"}, {"kind": "method", "line": 1277, "name": "eot_token", "signature": "def eot_token(self)"}, {"kind": "method", "line": 1389, "name": "__init__", "signature": "def __init__(self, root, cache_dir, logger)"}, {"doc": "Walk directory tree collecting text file paths. Cached to disk.", "kind": "method", "line": 1396, "name": "scan", "signature": "def scan(self, force)"}, {"kind": "method", "line": 1463, "name": "__init__", "signature": "def __init__(self, cache_dir, logger)"}, {"doc": "Tokenize all files and return a memory-mapped numpy array.\n\nArgs:\n    file_paths: List of absolute file paths to tokenize.\n    tokenizer: BPE tokenizer instance.\n    cache_key: Unique key for caching tokens to disk.\n    max_tokens: Maximum number of tokens to produce.\n    min_chars: Skip files with fewer characters.\n\nReturns:\n    np.ndarray backed by a memmap on disk. Only accessed pages\n    are loaded into RAM by the OS virtual memory system.", "kind": "method", "line": 1468, "name": "tokenize", "signature": "def tokenize(self, file_paths, tokenizer, cache_key, max_tokens, min_chars)"}, {"kind": "method", "line": 1551, "name": "__init__", "signature": "def __init__(self, tokens, seq_len)"}, {"kind": "method", "line": 1556, "name": "__len__", "signature": "def __len__(self)"}, {"kind": "method", "line": 1559, "name": "__getitem__", "signature": "def __getitem__(self, idx)"}, {"kind": "method", "line": 1579, "name": "__init__", "signature": "def __init__(self, config, logger)"}, {"doc": "Shannon entropy of byte frequencies (bits per byte).", "kind": "method", "line": 1588, "name": "_compute_entropy", "signature": "def _compute_entropy(self, text)"}, {"doc": "Return True if any line exceeds threshold characters.", "kind": "method", "line": 1602, "name": "_has_long_lines", "signature": "def _has_long_lines(self, text, threshold)"}, {"doc": "Fraction of tokens that are pure whitespace or indentation-only.", "kind": "method", "line": 1609, "name": "_special_token_ratio", "signature": "def _special_token_ratio(self, text, tokenizer)"}, {"kind": "method", "line": 1623, "name": "_content_hash", "signature": "def _content_hash(self, text)"}, {"doc": "Read and evaluate a file. Returns text if passed, None if filtered.", "kind": "method", "line": 1626, "name": "filter_file", "signature": "def filter_file(self, path, tokenizer)"}, {"kind": "method", "line": 1666, "name": "report", "signature": "def report(self)"}, {"kind": "method", "line": 1687, "name": "__init__", "signature": "def __init__(self, tokens, seq_len, file_tiers, active_tier, logger)"}, {"kind": "method", "line": 1697, "name": "_update_len", "signature": "def _update_len(self)"}, {"kind": "method", "line": 1703, "name": "set_tier", "signature": "def set_tier(self, tier)"}, {"kind": "method", "line": 1707, "name": "__len__", "signature": "def __len__(self)"}, {"kind": "method", "line": 1710, "name": "__getitem__", "signature": "def __getitem__(self, idx)"}, {"kind": "method", "line": 1755, "name": "__init__", "signature": "def __init__(self, base_trainer)"}, {"kind": "method", "line": 1761, "name": "_build_dataloader", "signature": "def _build_dataloader(self, dataset, seq_len, batch_size, is_train)"}, {"doc": "Run training with progressive sequence length across phases.", "kind": "method", "line": 1771, "name": "run", "signature": "def run(self, train_paths, val_paths, tokenizer, file_tiers, phases)"}, {"kind": "method", "line": 1829, "name": "__init__", "signature": "def __init__(self, target_model, config, logger)"}, {"kind": "method", "line": 1837, "name": "_build_draft", "signature": "def _build_draft(self)"}, {"doc": "Autoregressive generation via speculative decoding.\n\nEach round: draft generates K tokens, target verifies all K in\none O(1) forward pass (longest context), then samples the first\nrejection from the target.", "kind": "method", "line": 1851, "name": "generate", "signature": "def generate(self, token_ids, max_new_tokens, temperature, top_k, repetition_penalty)"}, {"kind": "method", "line": 1942, "name": "__init__", "signature": "def __init__(self, embed, mode)"}, {"kind": "method", "line": 1971, "name": "forward", "signature": "def forward(self, indices)"}, {"kind": "method", "line": 2011, "name": "__init__", "signature": "def __init__(self, model, config, tokenizer)"}, {"kind": "method", "line": 2018, "name": "cache_tokens", "signature": "def cache_tokens(self, key, tokens)"}, {"kind": "method", "line": 2022, "name": "model", "signature": "def model(self)"}, {"kind": "method", "line": 2026, "name": "optimizer", "signature": "def optimizer(self)"}, {"kind": "method", "line": 2030, "name": "scaler", "signature": "def scaler(self)"}, {"kind": "method", "line": 2034, "name": "amp_dtype", "signature": "def amp_dtype(self)"}, {"kind": "method", "line": 2038, "name": "completed_epochs", "signature": "def completed_epochs(self)"}, {"kind": "method", "line": 2042, "name": "completed_epochs", "signature": "def completed_epochs(self, v)"}, {"kind": "method", "line": 2046, "name": "global_step", "signature": "def global_step(self)"}, {"kind": "method", "line": 2050, "name": "global_step", "signature": "def global_step(self, v)"}, {"kind": "method", "line": 2054, "name": "best_val_loss", "signature": "def best_val_loss(self)"}, {"kind": "method", "line": 2058, "name": "best_val_loss", "signature": "def best_val_loss(self, v)"}, {"kind": "method", "line": 2062, "name": "history", "signature": "def history(self)"}, {"kind": "method", "line": 2066, "name": "ckpt_mgr", "signature": "def ckpt_mgr(self)"}, {"kind": "method", "line": 2069, "name": "resume", "signature": "def resume(self)"}, {"kind": "method", "line": 2072, "name": "_current_state", "signature": "def _current_state(self)"}, {"kind": "method", "line": 2075, "name": "_cosine_lr", "signature": "def _cosine_lr(self)"}, {"kind": "method", "line": 2078, "name": "_set_lr", "signature": "def _set_lr(self)"}, {"kind": "method", "line": 2081, "name": "evaluate", "signature": "def evaluate(self, dataloader)"}, {"kind": "method", "line": 2084, "name": "_sample_text", "signature": "def _sample_text(self)"}, {"doc": "Training loop with progressive sequence length across phases.", "kind": "method", "line": 2087, "name": "_progressive_train", "signature": "def _progressive_train(self, train_paths, val_paths, tokenizer, phases, memtok)"}, {"kind": "method", "line": 2126, "name": "train", "signature": "def train(self, train_dl, val_dl)"}, {"doc": "Top-level entry point: curriculum + progressive seq len.", "kind": "method", "line": 2129, "name": "run_curriculum", "signature": "def run_curriculum(self, train_paths, val_paths, tokenizer, phases)"}, {"kind": "method", "line": 2180, "name": "__init__", "signature": "def __init__(self, config, logger)"}, {"doc": "Lee el checkpoint 'latest' y ajusta cfg.N_KV_HEADS / cfg.GQA_GROUPS\npara que coincidan con la arquitectura guardada.\nNecesario cuando el codigo cambio GQA despues de guardar el checkpoint.", "kind": "method", "line": 2190, "name": "patch_config_for_resume", "signature": "def patch_config_for_resume(self, cfg)"}, {"kind": "method", "line": 2219, "name": "_save_model", "signature": "def _save_model(self, model, directory)"}, {"kind": "method", "line": 2232, "name": "_load_model", "signature": "def _load_model(self, model, directory)"}, {"kind": "method", "line": 2263, "name": "_save_optimizer", "signature": "def _save_optimizer(self, optimizer, directory)"}, {"kind": "method", "line": 2266, "name": "_load_optimizer", "signature": "def _load_optimizer(self, optimizer, directory, device)"}, {"kind": "method", "line": 2275, "name": "_save_state", "signature": "def _save_state(self, state, directory)"}, {"kind": "method", "line": 2280, "name": "_load_state", "signature": "def _load_state(self, directory)"}, {"kind": "method", "line": 2291, "name": "should_save", "signature": "def should_save(self)"}, {"doc": "Guarda checkpoint completo.\n\nstate debe contener al menos: completed_epochs, global_step,\nbest_val_loss, history, config.", "kind": "method", "line": 2294, "name": "save", "signature": "def save(self, model, optimizer, state, is_best)"}, {"doc": "Carga el ultimo checkpoint guardado.\nDevuelve el state dict (vacio si no hay checkpoint).", "kind": "method", "line": 2339, "name": "load_latest", "signature": "def load_latest(self, model, optimizer)"}, {"doc": "Carga el mejor modelo guardado (solo pesos, sin optimizador).", "kind": "method", "line": 2366, "name": "load_best", "signature": "def load_best(self, model)"}, {"kind": "method", "line": 2378, "name": "has_checkpoint", "signature": "def has_checkpoint(self)"}, {"kind": "method", "line": 2400, "name": "__init__", "signature": "def __init__(self, model, config, tokenizer)"}, {"doc": "Carga el ultimo checkpoint disponible.\nRestaura: pesos del modelo, estado del optimizador, historial acumulado,\nepoch/step completados y mejor val_loss.\nDevuelve True si se cargo un checkpoint, False si empieza de cero.", "kind": "method", "line": 2435, "name": "resume", "signature": "def resume(self)"}, {"doc": "Construye el dict de estado para persistir en state.json.", "kind": "method", "line": 2460, "name": "_current_state", "signature": "def _current_state(self)"}, {"doc": "Cosine decay con warmup. El schedule es relativo a la sesion actual.", "kind": "method", "line": 2471, "name": "_cosine_lr", "signature": "def _cosine_lr(self, step_in_session, total_steps_session)"}, {"kind": "method", "line": 2479, "name": "_set_lr", "signature": "def _set_lr(self, lr)"}, {"doc": "Entrena cfg.EPOCHS epocas adicionales a partir de completed_epochs.\nEl historial se acumula sobre sesiones previas.", "kind": "method", "line": 2483, "name": "train", "signature": "def train(self, train_dl, val_dl)"}, {"doc": "Genera una muestra de texto al final de cada epoch para monitorear\nla calidad cualitativa del modelo (detecta degeneracion, repeticion, etc.).", "kind": "method", "line": 2619, "name": "_sample_text", "signature": "def _sample_text(self, tokenizer, prompts, max_new, temperature, top_k)"}, {"kind": "method", "line": 2651, "name": "evaluate", "signature": "def evaluate(self, dataloader)"}, {"kind": "method", "line": 2701, "name": "__init__", "signature": "def __init__(self, config)"}, {"kind": "method", "line": 2709, "name": "compute_delta", "signature": "def compute_delta(self, model)"}, {"kind": "method", "line": 2716, "name": "compute_alpha", "signature": "def compute_alpha(self, delta)"}, {"doc": "Captura gradientes de forma segura, ignorando tensores corruptos.", "kind": "method", "line": 2721, "name": "update_grad_buffer", "signature": "def update_grad_buffer(self, model)"}, {"doc": "T_eff = lr/2 * Var(gradiente). Temperatura termodinamica efectiva.", "kind": "method", "line": 2747, "name": "compute_t_eff", "signature": "def compute_t_eff(self, lr)"}, {"doc": "κ = λ_max / λ_min de la covarianza del gradiente.\nParámetro de orden para cristalización (κ≈1 = cristal).\nNota: requiere pasadas backward adicionales. Se ejecuta con protección\npara no corromper el estado AMP del trainer principal.", "kind": "method", "line": 2755, "name": "compute_kappa", "signature": "def compute_kappa(self, model, dataloader, n_batches)"}, {"doc": "Fase de Berry de los kernels espectrales imaginarios.\nSurge de los parametros ki_w, ki_x, ki_y, ki_z de QuaternionSpectralLayer.\n|berry|>pi/2 con winding!=0 indica estructura topologica.", "kind": "method", "line": 2813, "name": "compute_berry_phase", "signature": "def compute_berry_phase(self, model)"}, {"doc": "Complejidad local: 1 - similitud coseno promedio entre filas de pesos.", "kind": "method", "line": 2826, "name": "compute_lc", "signature": "def compute_lc(self, model)"}, {"doc": "Superposicion: correlacion inter-fila promedio (entrelazamiento de features).", "kind": "method", "line": 2840, "name": "compute_sp", "signature": "def compute_sp(self, model)"}, {"doc": "Clasificacion de fase segun Book.md:\n\ndiscrete_crystal:       delta<0.05, kappa<1.5\ntopological_insulator:  |berry|>pi/2, winding!=0\ncold_glass:             kappa>>1, delta>0.3\nfunctional_glass:       intermedio (lo mas comun en LM)", "kind": "method", "line": 2856, "name": "classify_phase", "signature": "def classify_phase(self, delta, kappa, berry)"}, {"doc": "Calcula todas las metricas.\ncompute_kappa=True hace pasadas backward adicionales (caro, usar cada N epochs).", "kind": "method", "line": 2875, "name": "compute_all", "signature": "def compute_all(self, model, lr, dataloader, compute_kappa)"}, {"kind": "method", "line": 2900, "name": "format_log", "signature": "def format_log(self, m)"}, {"kind": "method", "line": 2936, "name": "__init__", "signature": "def __init__(self, config, logger)"}, {"doc": "Mide la coherencia espectral para un ratio dado.\nRetorna: varianza del gradiente (menor = mas coherente = mejor).", "kind": "method", "line": 2940, "name": "_measure_ratio", "signature": "def _measure_ratio(self, ratio, sample_batch)"}, {"doc": "Retorna el mejor ratio de inicializacion de kernels espectrales.", "kind": "method", "line": 2969, "name": "optimize", "signature": "def optimize(self, dataloader)"}, {"kind": "method", "line": 3009, "name": "__init__", "signature": "def __init__(self, config, logger)"}, {"doc": "Retorna el mejor batch size segun delta y T_eff.", "kind": "method", "line": 3013, "name": "prospect", "signature": "def prospect(self, candidates, train_dataset, prospect_steps)"}, {"kind": "method", "line": 3092, "name": "__init__", "signature": "def __init__(self, config, logger)"}, {"doc": "Retorna la semilla con la mejor trayectoria de delta.", "kind": "method", "line": 3096, "name": "mine", "signature": "def mine(self, seed_start, n_seeds, train_dataset, prospect_steps)"}, {"kind": "method", "line": 3178, "name": "__init__", "signature": "def __init__(self, trainer, t0, cooling_rate, stagnation_patience)"}, {"doc": "Ejecuta refine_epochs epocas de recocido simulado.\nRetorna el historial de refinamiento.", "kind": "method", "line": 3187, "name": "refine", "signature": "def refine(self, train_dl, val_dl, refine_epochs)"}, {"kind": "method", "line": 3330, "name": "__init__", "signature": "def __init__(self, config, train_tokens, val_tokens, tokenizer, logger, curriculum_tiers, progressive_seq)"}, {"kind": "method", "line": 3343, "name": "_build_dataloader", "signature": "def _build_dataloader(self, tokens, seq_len, batch_size, shuffle, tag)"}, {"kind": "method", "line": 3357, "name": "_build_phases", "signature": "def _build_phases(self)"}, {"kind": "method", "line": 3366, "name": "run", "signature": "def run(self, run_prospect, refine_epochs, resume, prospect_steps, probe_seeds, seed_start)"}, {"kind": "method", "line": 3468, "name": "__init__", "signature": "def __init__(self, config, train_dataset, val_dataset, tokenizer, logger)"}, {"kind": "method", "line": 3478, "name": "_make_dataloaders", "signature": "def _make_dataloaders(self, batch_size)"}, {"doc": "Ejecuta el pipeline completo.\nRetorna el trainer con el modelo entrenado.", "kind": "method", "line": 3490, "name": "run", "signature": "def run(self, run_prospect, refine_epochs, resume, prospect_steps, probe_seeds, seed_start)"}, {"kind": "method", "line": 1031, "name": "ckpt_fn", "signature": "def ckpt_fn(x_in)"}]}, {"doc": "TopoGPT3: Grassmannian / Berry-Holonomy extension of TopoGPT2  Author: Gris Iscomeback License: GPL v3  Lo nuevo respecto a model.py --------------------------------- 1. Espacio base: Grassmanniana Gr(r, N) sobre el tensor de kernels espectrales K(theta) en C^{N_f x N_c}. El estado geometrico vive en U_r(theta) en St(r,N)/U(r), con r elegido dinamicamente por el \"elbow\" del espectro singular de K. 2. Fisher gap funcional:    Delta_F(theta) = lambda_r(Sigma_F) - lambda_{r+1}(Sigma_F) estimado por covarianza empirica de gradientes (mini-batch) o por scores. 3. Conexion de Berry discreta:  A_n = i * U_n^dagger (U_{n+1} - U_n) y holonomia acumulada      U_Gamma = P prod_n exp(-i A_n)  en U(r). 4. Distancia de conjugacion en SU(2) (cuaternionico, r=1 efectivo): d_conj(U1, U2) = min_{g in SU(2)} || U1 - g U2 g^{-1} ||_F 5. Winding W como proxy heuristico barato (rol secundario). 6. Curriculum por dataset, de mas simple a mas dificil: Tier 1: CodeAlpaca               (instrucciones cortas) Tier 2: Code-Feedback-Filtered   (chat / explicacion paso a paso) Tier 3: Magicoder-Evol-Instruct-110K (problemas complejos) Tier 4: Tiny-The-Stack           (codigo real multilenguaje) Cada tier mantiene splits train / val / holdout *disjuntos*. El conjunto HOLDOUT nunca se ve durante entrenamiento; se usa solo para medir generalizacion verdadera al final de cada tier y al final del pipeline.  Diseno", "id": "topogpt3/train.py", "kind": "module", "label": "train.py", "language": "py", "sha256": "12be4427fea7b715", "symbol_count": 64, "symbols": [{"doc": "Configuracion del pipeline TopoGPT3 (Grassmanniana + curriculum).", "kind": "class", "line": 87, "name": "TopoGPT3Config", "signature": "class TopoGPT3Config"}, {"doc": "Observables geometricos sobre la trayectoria SGD.\n\nEn cada snapshot:\n  - Apila los kernels espectrales (kr_*, ki_*) del modelo en\n    K(theta) en C^{N_f x N_c}.\n  - SVD truncada -> U_r(theta) en St(r,N).\n  - Rango r dinamico por elbow de los valores singulares.\n  - Gap funcional Delta_F estimado por covarianza de gradientes\n    muestrales (proxy de la matriz de Fisher).\n  - Conexion de Berry discreta entre snapshots consecutivos:\n       A_n = i * U_n^dagger (U_{n+1} - U_n)\n    Holonomia acumulada U_Gamma = P prod_n exp(-i A_n) en U(r).\n  - Distancia de conjugacion en SU(2) (r=1 efectivo cuaternionico).\n  - Winding W como proxy barato.\n\nTodos los calculos viven en CPU/float32 para no contaminar AMP.", "kind": "class", "line": 209, "name": "GrassmannianTracker", "signature": "class GrassmannianTracker"}, {"doc": "Sustituye QuaternionSpectralLayer._contract usando el truco de Gauss.\n\nPara (Wr + i Wi)(Xr + i Xi) la version naive requiere 4 productos reales:\n    Yr = Wr Xr - Wi Xi\n    Yi = Wr Xi + Wi Xr\nGauss (Karatsuba) baja a 3 productos reales:\n    m1 = Wr * Xr\n    m2 = Wi * Xi\n    m3 = (Wr + Wi) * (Xr + Xi)\n    Yr = m1 - m2\n    Yi = m3 - m1 - m2\n\nImportante (AMP): el _contract original opera sobre complex64 y PyTorch no\nautocastea operaciones complejas; el resultado es complex64. Si dejamos que\nautocast convierta nuestros einsums reales a fp16, la dtype de salida cambia\ny rompe el scatter_add_ corriente abajo en QuaternionTorusBrain. Por eso\ndesactivamos autocast aqui y forzamos fp32 para preservar la semantica.", "kind": "method", "line": 544, "name": "_gauss_complex_contract", "signature": "def _gauss_complex_contract(self, W, X)"}, {"doc": "Activa la version Gauss de _contract en QuaternionSpectralLayer.\nIdempotente: solo parchea una vez por proceso.", "kind": "method", "line": 580, "name": "apply_gauss_patch", "signature": "def apply_gauss_patch(logger)"}, {"doc": "Mide y calcula los tres ratios pedidos:\n\n  perf_per_param  =  (1 / val_ppl) / params_M\n  perf_per_FLOP   =  tokens_per_sec / FLOPs_per_sec_aprox\n  perf_per_BW     =  tokens_per_sec / bytes_moved_per_sec_aprox\n\nFLOPs estimados con la heuristica de Kaplan/Hoffmann:\n    FLOPs_forward_per_token ~= 2 * N_no_embed\n    FLOPs_total_per_token  ~= 6 * N_no_embed       (forward + backward)\nBandwidth estimada como params_bytes leidos + activations_bytes movidas por step.\ntokens_per_sec se cronometra empiricamente sobre el dataloader.", "kind": "class", "line": 597, "name": "EfficiencyMetrics", "signature": "class EfficiencyMetrics"}, {"doc": "    Carga los 4 datasets, normaliza cada ejemplo a una unica cadena de texto,\n    tokeniza con BPE y produce splits train / val / holdout disjuntos.\n\n    Politica de normalizacion por dataset:\n      - CodeAlpaca:           \"### Instruction\n{i}\n### Input\n{x}\n### Response\n{o}\"\n      - Code-Feedback:        concat de turnos: \"<usr> ... </usr>\n<asst> ... </asst>\"\n      - Magicoder-Evol:       \"### Problem\n{p}\n### Solution\n{s}\"\n      - Tiny-The-Stack:       texto crudo del archivo (truncado a 32k chars/file)\n\n    Cache en disco: tokens_{tier}_{split}.bin (int32 memmap) + manifest .json.\n    El HOLDOUT se separa con seed fija antes de tokenizar para garantizar\n    que la misma muestra nunca aparezca en train o val entre corridas.\n    ", "kind": "class", "line": 725, "name": "CodeCurriculumLoader", "signature": "class CodeCurriculumLoader"}, {"doc": "Dataset autoregresivo sobre un stream de tokens.\nCada item es (x, y) con shape [seq_len].", "kind": "class", "line": 1012, "name": "BlockTokenDataset", "signature": "class BlockTokenDataset(Dataset)"}, {"doc": "Persiste pesos del modelo + estado del trainer (sin AMP scaler para portabilidad).", "kind": "class", "line": 1039, "name": "CheckpointStore", "signature": "class CheckpointStore"}, {"doc": "Orquesta el curriculum sobre los 4 tiers.\n\nPipeline por tier:\n  1. Abre memmap de tokens (train/val/holdout).\n  2. Construye DataLoaders con seq_len(tier).\n  3. Entrena TIER_EPOCHS[tier] epocas con AMP + grad accum.\n  4. Cada GRASS_TRACK_EVERY steps: snapshot Grassmanniano.\n  5. Al final de cada epoca: eval en VAL.\n  6. Al final del tier: eval en HOLDOUT (datos nunca vistos).\n  7. Checkpoint y avanza al siguiente tier.\n\nAl final del pipeline: eval en HOLDOUT *combinado* de los 4 tiers.", "kind": "class", "line": 1115, "name": "TopoGPT3Trainer", "signature": "class TopoGPT3Trainer"}, {"kind": "method", "line": 1728, "name": "parse_args", "signature": "def parse_args()"}, {"kind": "method", "line": 1764, "name": "main", "signature": "def main()"}, {"kind": "method", "line": 182, "name": "build_topogpt2_config", "signature": "def build_topogpt2_config(self, max_seq_len, attn_window)"}, {"kind": "method", "line": 229, "name": "__init__", "signature": "def __init__(self, config, logger)"}, {"doc": "Devuelve K(theta) en C^{N_f x N_c}:\n  - filas = frecuencias planas (todos los modos espaciales de todos los kernels)\n  - columnas = canales (in_q * out_q por componente cuaternionico, sumados)", "kind": "method", "line": 243, "name": "_stack_spectral_kernels", "signature": "def _stack_spectral_kernels(model)"}, {"doc": "Punto donde el valor singular cae por debajo de elbow_ratio * sigma_max.", "kind": "method", "line": 280, "name": "_elbow_rank", "signature": "def _elbow_rank(self, sigmas)"}, {"doc": "SVD compacta y truncada.\nDevuelve (U_r, sigmas, r) con U_r en C^{N_f x r} ortonormal.", "kind": "method", "line": 289, "name": "_dominant_subspace", "signature": "def _dominant_subspace(self, K)"}, {"doc": "Concatena un sub-sample de gradientes para mantener costo acotado.", "kind": "method", "line": 307, "name": "_flatten_grads", "signature": "def _flatten_grads(model, max_per_tensor)"}, {"doc": "Sigma_F ~= (1/M) sum_m g_m g_m^T  (covarianza muestral de gradientes).\nDelta_F = lambda_{r_eff} - lambda_{r_eff+1}, donde r_eff = min(r_target, M-2)\npara no salir del rango efectivo del estimador con M gradientes.\nDevuelve (gap, eigs_desc, r_eff).", "kind": "method", "line": 325, "name": "estimate_fisher_gap", "signature": "def estimate_fisher_gap(self, model, dataloader, vocab_size, r_target)"}, {"doc": "Proyeccion a U(r) por descomposicion polar (M ~= U H -> retorna U).", "kind": "method", "line": 393, "name": "_project_unitary", "signature": "def _project_unitary(M)"}, {"doc": "Holonomia discreta:\n    T_n = U_n^dagger U_{n+1}  en C^{r x r}  (transporte paralelo discreto)\n    U_Gamma <- T_n * U_Gamma  (acumulado)\nTras cada paso, U_Gamma se proyecta a U(r) para evitar deriva numerica.", "kind": "method", "line": 398, "name": "update_holonomy", "signature": "def update_holonomy(self, U_new)"}, {"doc": "Para U1, U2 en U(1)/U(2):  d_conj(U1, U2) = min_g || U1 - g U2 g^{-1} ||_F.\nEn U(1) coincide con |U1 - U2|.\nEn SU(2) se reduce a comparar |Tr(U1)| con |Tr(U2)| (clase de conjugacion).", "kind": "method", "line": 424, "name": "conjugation_distance_su2", "signature": "def conjugation_distance_su2(U1, U2)"}, {"doc": "W += (1/2pi) * arg det <U_prev | U_new>  acumulado sobre la trayectoria.", "kind": "method", "line": 441, "name": "_accumulate_winding", "signature": "def _accumulate_winding(self, U_new)"}, {"kind": "method", "line": 457, "name": "snapshot", "signature": "def snapshot(self, model, step, dataloader, vocab_size)"}, {"kind": "method", "line": 512, "name": "format_log", "signature": "def format_log(self, snap)"}, {"kind": "method", "line": 534, "name": "save", "signature": "def save(self, path)"}, {"kind": "method", "line": 612, "name": "__init__", "signature": "def __init__(self, model, config, logger, gauss_enabled)"}, {"kind": "method", "line": 623, "name": "_embed_params", "signature": "def _embed_params(model)"}, {"doc": "Devuelve (tokens_por_segundo, segundos_por_step).", "kind": "method", "line": 631, "name": "measure_throughput", "signature": "def measure_throughput(self, dataloader, vocab_size)"}, {"doc": "Heuristica: 6 * N_no_embed * tokens (forward + backward).", "kind": "method", "line": 663, "name": "estimate_flops_per_step", "signature": "def estimate_flops_per_step(self, batch_size, seq_len)"}, {"doc": "Bandwidth aproximada: lectura de pesos + activaciones por step.\nAsume AMP fp16 (2 bytes); pesos fp32 (4 bytes) leidos una vez.", "kind": "method", "line": 668, "name": "estimate_bytes_per_step", "signature": "def estimate_bytes_per_step(self, batch_size, seq_len, dtype_bytes)"}, {"kind": "method", "line": 676, "name": "compute", "signature": "def compute(self, dataloader, vocab_size, val_loss, val_ppl, val_acc, batch_size, seq_len)"}, {"kind": "method", "line": 708, "name": "format_log", "signature": "def format_log(self, m)"}, {"kind": "method", "line": 741, "name": "__init__", "signature": "def __init__(self, config, tokenizer, logger)"}, {"kind": "method", "line": 758, "name": "_format_codealpaca", "signature": "def _format_codealpaca(ex)"}, {"kind": "method", "line": 769, "name": "_format_code_feedback", "signature": "def _format_code_feedback(ex)"}, {"kind": "method", "line": 789, "name": "_format_magicoder", "signature": "def _format_magicoder(ex)"}, {"kind": "method", "line": 797, "name": "_format_tiny_stack", "signature": "def _format_tiny_stack(ex)"}, {"kind": "method", "line": 809, "name": "_get_formatter", "signature": "def _get_formatter(cls, tier)"}, {"kind": "method", "line": 831, "name": "_tier_paths", "signature": "def _tier_paths(self, tier)"}, {"kind": "method", "line": 837, "name": "_manifest_path", "signature": "def _manifest_path(self, tier)"}, {"doc": "True solo si los 3 splits existen, son no-vacios y el manifest concuerda.", "kind": "method", "line": 840, "name": "_already_prepared", "signature": "def _already_prepared(self, tier)"}, {"doc": "Carga el dataset HF; para tiny_the_stack prueba una cadena de fallbacks\npublicos hasta que uno funcione.", "kind": "method", "line": 870, "name": "_load_hf_with_fallback", "signature": "def _load_hf_with_fallback(self, tier)"}, {"kind": "method", "line": 903, "name": "prepare_tier", "signature": "def prepare_tier(self, tier_index, force)"}, {"kind": "method", "line": 998, "name": "open_memmap", "signature": "def open_memmap(self, tier, split)"}, {"kind": "method", "line": 1018, "name": "__init__", "signature": "def __init__(self, tokens, seq_len)"}, {"kind": "method", "line": 1023, "name": "__len__", "signature": "def __len__(self)"}, {"kind": "method", "line": 1026, "name": "__getitem__", "signature": "def __getitem__(self, idx)"}, {"kind": "method", "line": 1042, "name": "__init__", "signature": "def __init__(self, root, max_keep, logger)"}, {"doc": "Guarda checkpoint atomico en <root>/last/ sobreescribiendo el anterior.\n\nEl argumento `tag` se conserva por compatibilidad pero se ignora: solo\nexiste un checkpoint llamado `last` y los pesos en safetensors.", "kind": "method", "line": 1049, "name": "save", "signature": "def save(self, tag, model, optimizer, state)"}, {"kind": "method", "line": 1083, "name": "load_latest", "signature": "def load_latest(self, model, optimizer)"}, {"kind": "method", "line": 1107, "name": "should_save", "signature": "def should_save(self, interval_min)"}, {"kind": "method", "line": 1131, "name": "__init__", "signature": "def __init__(self, config, start_tier, exploitgym_loader, topogpt3_loader, merged_config)"}, {"kind": "method", "line": 1203, "name": "prepare_all", "signature": "def prepare_all(self, force)"}, {"kind": "method", "line": 1244, "name": "_build_loaders", "signature": "def _build_loaders(self, tier_index)"}, {"kind": "method", "line": 1294, "name": "_cosine_lr", "signature": "def _cosine_lr(self, step, total_steps)"}, {"kind": "method", "line": 1301, "name": "_set_lr", "signature": "def _set_lr(self, lr)"}, {"kind": "method", "line": 1309, "name": "_train_one_tier", "signature": "def _train_one_tier(self, tier_index)"}, {"doc": "Evaluate holdout of ALL previous tiers to detect catastrophic forgetting.", "kind": "method", "line": 1527, "name": "_check_forgetting", "signature": "def _check_forgetting(self, completed_tier_index)"}, {"doc": "Open train/val/holdout memmaps for a given tier.", "kind": "method", "line": 1557, "name": "_open_memmaps", "signature": "def _open_memmaps(self, tier_index)"}, {"doc": "Devuelve (avg_loss, perplexity, token_accuracy).", "kind": "method", "line": 1585, "name": "_evaluate", "signature": "def _evaluate(self, dl)"}, {"kind": "method", "line": 1623, "name": "_state_dict", "signature": "def _state_dict(self)"}, {"kind": "method", "line": 1636, "name": "run", "signature": "def run(self)"}, {"kind": "method", "line": 1693, "name": "_eval_combined_holdout", "signature": "def _eval_combined_holdout(self)"}, {"kind": "method", "line": 934, "name": "flush", "signature": "def flush(split)"}]}, {"doc": "Vanilla transformer control for TopoGPT3 Hodge-CM ablation. Matched size: D=384, L=8, H=8, SwiGLU FFN, RoPE, RMSNorm, tied embeddings. Target: ~33M to compare against TopoGPT3 37.4M (which includes TorusBrain+MoE+spectral). Same vocab (50257), full curriculum interface (codealpaca, code_feedback, magicoder_evol, tiny_the_stack) with train/val/holdout disjuntos — nada recortado. Usa tokens reales del tier cuando existen; sintetico solo con --allow-synthetic explícito.", "id": "topogpt3/vanilla_control.py", "kind": "module", "label": "vanilla_control.py", "language": "py", "sha256": "479e674d2779ee38", "symbol_count": 26, "symbols": [{"kind": "function", "line": 22, "name": "resolve_device", "signature": "def resolve_device(req)"}, {"doc": "Busca train real tokenizado del tier. Devuelve lista de ints o None.\nNada silencioso: si no hay, el llamador decide (error o sintetico explícito).", "kind": "function", "line": 41, "name": "find_tier_tokens", "signature": "def find_tier_tokens(tier)"}, {"kind": "class", "line": 61, "name": "VanillaConfig", "signature": "class VanillaConfig"}, {"kind": "class", "line": 72, "name": "RMSNorm", "signature": "class RMSNorm(Module)"}, {"kind": "class", "line": 83, "name": "RotaryEmbedding", "signature": "class RotaryEmbedding(Module)"}, {"kind": "class", "line": 113, "name": "CausalAttention", "signature": "class CausalAttention(Module)"}, {"kind": "class", "line": 134, "name": "SwiGLU", "signature": "class SwiGLU(Module)"}, {"kind": "class", "line": 146, "name": "VanillaLayer", "signature": "class VanillaLayer(Module)"}, {"doc": "Tied-embedding decoder-only transformer. Same training interface as TopoGPT2.", "kind": "class", "line": 161, "name": "VanillaTransformer", "signature": "class VanillaTransformer(Module)"}, {"kind": "method", "line": 191, "name": "build_optimizer", "signature": "def build_optimizer(model, cfg)"}, {"kind": "method", "line": 73, "name": "__init__", "signature": "def __init__(self, d, eps)"}, {"kind": "method", "line": 78, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 84, "name": "__init__", "signature": "def __init__(self, d_head, max_seq_len, base)"}, {"kind": "method", "line": 90, "name": "_build_cache", "signature": "def _build_cache(self, seq_len)"}, {"kind": "method", "line": 98, "name": "_rot_half", "signature": "def _rot_half(x)"}, {"kind": "method", "line": 102, "name": "forward", "signature": "def forward(self, q, k)"}, {"kind": "method", "line": 114, "name": "__init__", "signature": "def __init__(self, cfg)"}, {"kind": "method", "line": 124, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 135, "name": "__init__", "signature": "def __init__(self, d)"}, {"kind": "method", "line": 142, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 147, "name": "__init__", "signature": "def __init__(self, cfg)"}, {"kind": "method", "line": 155, "name": "forward", "signature": "def forward(self, x)"}, {"kind": "method", "line": 164, "name": "__init__", "signature": "def __init__(self, cfg)"}, {"kind": "method", "line": 176, "name": "_init", "signature": "def _init(m)"}, {"kind": "method", "line": 180, "name": "forward", "signature": "def forward(self, ids)"}, {"kind": "method", "line": 186, "name": "count_params", "signature": "def count_params(self)"}]}, {"doc": "Transfer weights from a smaller TopoGPT model to a larger one.  Strategy: - Matching layers (same shape): direct copy - Embedding/lm_head (vocab same, dim larger): pad with zeros - Linear layers (in or out dim larger): copy sub-block, pad rest - QuaternionSpectralLayer kernels: copy matching freq bins, pad rest - New layers (no match): random init (default torch init)  This gives the larger model a head start from the smaller model's knowledge.  Usage: python transfer_weights.py --from checkpoints_topogpt3/last/model.safetensors \\\\ --to checkpoints_topoexploit/last/model.safetensors \\\\ --scale-from small --scale-to large", "id": "transfer_weights.py", "kind": "module", "label": "transfer_weights.py", "language": "py", "sha256": "704749b6999f14a4", "symbol_count": 3, "symbols": [{"doc": "Transfer weights from small model state_dict to large model.\nReturns the new state_dict for the large model.", "kind": "function", "line": 31, "name": "transfer_weights", "signature": "def transfer_weights(state_small, model_large, logger)"}, {"doc": "Copy small tensor into the top-left of a larger tensor, rest stays as large's init.", "kind": "function", "line": 135, "name": "_copy_with_padding", "signature": "def _copy_with_padding(small, large)"}, {"kind": "function", "line": 147, "name": "main", "signature": "def main()"}]}], "type": "CodePropertyGraph", "version": "1.0"}
```

---

## Architecture Reference

### PY (39 files)

#### `app.py`
**Path:** `app.py`
**File Doc:** *Drop-in entry point that demonstrates how to use the topogpt3 package.  This file lives outside the package on purpose. Copy it (or its sections) into your own project after running ``pip install topogpt3``. Three usage patterns are shown:  1. ``run_inference`` calls the standard autoregressive sampler. 2. ``run_inference_hrm`` calls the hierarchical recursive reasoning sampler that reuses the same checkpoint with no extra trained parameters. 3. ``run_training`` launches the full curriculum trainer.  The script's main() exposes them through a tiny ``--mode`` CLI so the file is runnable as-is for a quick smoke test once a checkpoint exists.*

**Functions:**
- `run_inference` (line 46) `def run_inference(prompt, checkpoint_dir, checkpoint_name, max_new_tokens, temperature, top_k, repetition_penalty, device)` - *Run the standard sampler and return the generated completion text.*
- `run_inference_hrm` (line 71) `def run_inference_hrm(prompt, checkpoint_dir, checkpoint_name, max_new_tokens, temperature, top_k, repetition_penalty, high_level_iters, low_level_iters, low_level_window, device)` - *Run the hierarchical recursive sampler and return the completion.*
- `run_training` (line 105) `def run_training(scale, start_tier, device, prepare_data)` - *Run the full TopoGPT3 curriculum trainer.*
- `_build_parser` (line 121) `def _build_parser()` - *Build the top-level CLI for this entry point script.*
- `main` (line 159) `def main(argv)` - *Entry point invoked when the file is executed as a script.*

#### `analyze.py`
**Path:** `eval/analyze.py`
**File Doc:** *Aggregate HumanEval result JSONL files into a summary table.  Reads one or more .jsonl files produced by harness.py and computes: - pass@1, pass@k (using the unbiased estimator from the HumanEval paper when k > 1) - mean latency, mean generation length, tok/s - per-error classification*

**Functions:**
- `pass_at_k` (line 21) `def pass_at_k(n, c, k)` - *Unbiased estimator from the HumanEval paper.

pass@k = 1 - C(n-c, k) / C(n, k)   if n - c >= k else 1.0
n = total samples, c = correct samples, k = target*
- `classify_error` (line 32) `def classify_error(msg, candidate_src)` - *Heuristic single-label error classifier.*
- `load_jsonl` (line 56) `def load_jsonl(path)`
- `summarize` (line 61) `def summarize(paths)`
- `main` (line 103) `def main()`

#### `analyze_results.py`
**Path:** `eval/analyze_results.py`
**File Doc:** *Analyze a HumanEval JSONL produced by harness.py.  For each failed problem the report shows: - the prompt fed to the model - the generated candidate after extraction - the hidden test that failed - the captured stdout/stderr and traceback  This makes it easy to see *how* and *why* a candidate failed without re-running the harness.  Usage: python eval/analyze_results.py eval/runs/run.jsonl python eval/analyze_results.py eval/runs/run.jsonl --summary python eval/analyze_results.py eval/runs/run.jsonl --task-id HumanEval/0*

**Functions:**
- `load_records` (line 26) `def load_records(path)`
- `summarize` (line 31) `def summarize(records)`
- `show_failures` (line 44) `def show_failures(records, task_id)`
- `main` (line 82) `def main()`

#### `diag_static.py`
**Path:** `eval/diag_static.py`
**File Doc:** *Diagnostico estatico de un checkpoint TopoGPT3 congelado.  Calcula sobre los pesos espectrales congelados (sin reentrenar):  kappa_F   = sigma_max / sigma_min  del kernel espectral apilado (proxy del condition number de la Grassmanniana) delta     = max |theta - round(theta)|  sobre los arg det de overlaps (cuantifica cuanto se "discretizan" las fases complejas) W         = (1/2pi) sum arg det <U_n | U_{n+1}>  (winding acumulado sobre barridos en frecuencia — sin trayectoria temporal real, usamos un barrido sintetico sobre los modos FFT) r         = rango dominante por elbow de los valores singulares sigma_*   = valores singulares principales  NOTA IMPORTANTE: Este script NO reentrena. Trabaja unicamente con los kernels espectrales cuaternionicos ya aprendidos. La "trayectoria" W se define barriendo sobre los modos de frecuencia (no sobre pasos de entrenamiento), asi que W aqui mide coherencia de fase intra-modelo, no winding temporal. Esta distincion se reporta explicitamente en el JSONL de salida.  Salida: eval/runs/diag_static_<timestamp>.jsonl*

**Functions:**
- `phase_discretization` (line 49) `def phase_discretization(K, n_samples, seed)` - *Muestrea n_samples overlaps aleatorios <u_i | u_j> sobre los vectores
singulares de K y mide cuanto se aleja su fase arg del reticulo 2*pi*Z.

delta = max |theta/2pi - round(theta/2pi)| sobre la muestra.

Tambien devuelve:
  delta_mean, delta_median, frac_near_integer (|.| < 0.05)*
- `synthetic_winding` (line 95) `def synthetic_winding(K, n_windows, window_size)` - *Como el checkpoint es estatico, no hay trayectoria temporal.
Construimos una pseudo-trayectoria deslizando una ventana sobre
los modos de frecuencia (filas de K) y acumulando arg det del
overlap entre ventanas consecutivas.

W = (1/2pi) sum_n arg det <U_{n} | U_{n+1}>*
- `static_kappa` (line 144) `def static_kappa(K)`
- `context_length_diagnostic` (line 171) `def context_length_diagnostic(model, tracker, device, lengths)`
- `main` (line 248) `def main()`

#### `governor.py`
**Path:** `eval/governor.py`
**File Doc:** *Streaming + governance for autoregressive generation.  Two classes that fix two real problems with the existing `topogpt3.inference` pipeline:  - `TokenStream` — a thread-safe queue that captures raw token IDs as they are produced by the model. Enables post-hoc prefix agreement and exact-match metrics that need the *raw* token stream (the current harness only stores the post-extracted candidate text, losing that information).  - `GenerationGovernor` — wraps `model.generate` and exposes stop hooks: per-token timeout, loop detection (last K tokens repeat), and a user-callable cancel. Returns a `GenerationResult` with the stop reason so callers can distinguish "ran out of tokens" from "hit the safety hook" from "user aborted".  This is a Python port of the patterns in `claude-code-main/src/utils/stream.ts` (Stream<T> AsyncIterator wrapper) and `claude-code-main/src/query/stopHooks.ts` (AsyncGenerator with `preventContinuation`). The TypeScript originals are 76 and 473 lines respectively; this module is ~150 lines because Python's GIL lets us avoid the manual promise queueing.  NOTE: This module does NOT modify `topogpt3/model.py`. The generation loop is replicated here (not monkey-patched) so the original `generate` remains the single source of truth for the production sampler.*

**Classes:**
- `TokenStream` (line 45) `class TokenStream` - *Thread-safe single-producer / single-consumer queue of token IDs.

The producer (the generation loop) calls `put(tok)` for each new
token. Consumers can iterate via `iter_tokens(block=True)` or
`drain()` to get everything emitted so far.

The stream tracks a monotonic counter so consumers can detect
"no new tokens since last call" cheaply.*
- `StopReason` (line 99) `class StopReason(str, Enum)`
- `GenerationResult` (line 109) `class GenerationResult` - *Outcome of a governed generation.*
- `GenerationGovernor` (line 134) `class GenerationGovernor` - *Run a model's autoregressive generation loop with optional stop
hooks and a streaming interface.

Usage:
    ts = TokenStream()
    governor = GenerationGovernor(
        model=model,
        ctx=prompt_tensor,
        stream=ts,
        max_new_tokens=256,
        temperature=0.2,
        top_k=40,
        repetition_penalty=1.1,
    )
    result = governor.run(stop_hooks=[loop_detector, timeout_hook])
    if result.stop_reason == StopReason.LOOP:
        ...*

**Methods:**
- `make_loop_detector` (line 285) `def make_loop_detector(window, min_repeats)` - *Return True if the last `window` tokens contain a sub-sequence
of length >= `min_repeats` that repeats consecutively.

Catches the "model is stuck in a loop" pathology where a 24M-param
model emits the same 4-token pattern indefinitely.*
- `make_timeout_hook` (line 314) `def make_timeout_hook(per_token_s)` - *Return True if the per-token wall time exceeds `per_token_s`.
Useful for catching token-generation stalls (rare on CPU, but
happens under memory pressure).*
- `__init__` (line 56) `def __init__(self)`
- `put` (line 62) `def put(self, tok)`
- `mark_done` (line 67) `def mark_done(self)`
- `drain` (line 72) `def drain(self)` - *Return all tokens emitted so far, atomic snapshot.*
- `wait_for_new` (line 77) `def wait_for_new(self, timeout)` - *Block up to `timeout` seconds for a new token. Returns True
if a new token arrived (or stream closed), False on timeout.*
- `is_closed` (line 86) `def is_closed(self)`
- `__len__` (line 90) `def __len__(self)`
- `__post_init__` (line 117) `def __post_init__(self)`
- `__init__` (line 156) `def __init__(self, model, ctx, stream, max_new_tokens, temperature, top_k, repetition_penalty, max_seq_len)`
- `cancel` (line 177) `def cancel(self)` - *Asynchronously stop the generation. Safe to call from any
thread (e.g. a watchdog thread or the main UI loop).*
- `_should_cancel` (line 182) `def _should_cancel(self)`
- `run` (line 185) `def run(self, stop_hooks)` - *Execute the generation loop. Returns when the model emits
EOS, hits max_new_tokens, a hook returns True, or cancel() is
called.*
- `hook` (line 292) `def hook(generated)`
- `hook` (line 320) `def hook(generated)`

#### `governor_smoke.py`
**Path:** `eval/governor_smoke.py`
**File Doc:** *Smoke test for eval.governor (TokenStream + GenerationGovernor).  Verifies: - TokenStream threadsafety with a producer/consumer scenario - GenerationGovernor emits one StopReason per call - Loop detector actually fires - User cancel() works*

**Functions:**
- `load_model` (line 30) `def load_model()`
- `test_tokenstream_threadsafety` (line 49) `def test_tokenstream_threadsafety()`
- `test_governor_basic` (line 79) `def test_governor_basic()`
- `test_loop_detector` (line 98) `def test_loop_detector()`
- `test_cancel` (line 118) `def test_cancel()`
- `producer` (line 53) `def producer()`
- `consumer` (line 59) `def consumer()`

#### `harness.py`
**Path:** `eval/harness.py`
**File Doc:** *Harness for evaluating TopoGPT3 on HumanEval (164 problems).  Faithful to the official HumanEval protocol: for each problem we feed the model the function signature and docstring, let it produce a completion, extract the candidate function (everything from `def` up to a sentinel), and run the hidden test against it. We do NOT use `entry_point` from the dataset because the prompt we feed the model already contains it.  Two sampler modes are supported: - "standard" -> topogpt3.InferencePipeline - "hrm"      -> topogpt3.HRMInferencePipeline  Results are written to JSONL so multiple sampler configurations can share a single HumanEval cache and be compared later.*

**Classes:**
- `ModelLoader` (line 217) `class ModelLoader` - *Build the model and tokenizer once, run many generations.*

**Functions:**
- `load_humaneval` (line 59) `def load_humaneval(cache_dir)`
- `build_prompt` (line 75) `def build_prompt(problem)` - *Return the exact prompt text fed to the model.

HumanEval's `prompt` field already contains the function signature and
docstring, with the body to be completed starting on the next line.*
- `extract_candidate` (line 100) `def extract_candidate(prompt, completion)` - *Combine prompt + completion into a single Python source string.

The completion may itself start with whitespace/indentation that
belongs inside the function body. We strip leading blank lines and
then concatenate; we also stop at the first top-level `def ` or
`class ` to avoid the model continuing with extra functions.

Robustness fixes:
  - Strip the special <|endoftext|> (GPT-2 EOT) token that the model
    emits at the end of every generation. Leaving it in the candidate
    produces a SyntaxError and zeroes the pass rate.
  - Drop any training-format delimiters (### Response, <|assistant|>,
    <|user|>) that leak from the instruction-tuning corpus.
  - Cut at the first top-level def/class/__main__ guard after the
    function body has started.*
- `run_one_test` (line 150) `def run_one_test(problem, candidate_src, timeout)` - *Execute the candidate against the hidden test.

Returns (passed, message, stdout, stderr, traceback). We follow HumanEval's
`evaluate` function: build namespace, exec the candidate, exec the test,
expect `check(candidate) == None`.*
- `run_one_test_sandboxed` (line 172) `def run_one_test_sandboxed(problem, candidate_src, timeout, sandbox_cfg)` - *Sandboxed variant of `run_one_test`. Runs the candidate in a
subprocess with stripped builtins, AST pre-check, and OS-enforced
timeout. Drop-in replacement: same 5-tuple return.

Enable by passing `--sandbox` to `harness.py` (not yet wired) or
by calling this function directly from your own evaluation script.*
- `make_sampler` (line 195) `def make_sampler(mode, settings_kwargs)` - *Backwards-compatible shim. The real implementation lives in
`eval.samplers` as a decorator-based registry. We re-export here
so existing imports of `from eval.harness import make_sampler`
keep working. New code should import from `eval.samplers`.*
- `completion_for_problem` (line 204) `def completion_for_problem(sampler, prompt)` - *Run a single completion and return (raw_output_text, metrics_dict).*

**Methods:**
- `evaluate_problem` (line 272) `def evaluate_problem(problem, loader, args, sample_idx)`
- `main` (line 315) `def main()`
- `__init__` (line 220) `def __init__(self, ckpt_dir, ckpt_name, device)`
- `generate` (line 246) `def generate(self, prompt, max_new_tokens, temperature, top_k, repetition_penalty)`

#### `hodge_cm_ablation.py`
**Path:** `eval/hodge_cm_ablation.py`
**File Doc:** *Hodge-CM offline ablation on real TopoGPT3 checkpoint. Runs on GPU when available (--device cuda), else CPU. No datasets removed: uses the real checkpoint kernels (all 12) + real layer weights; nothing subsampled, no tier dropped. Saves JSON to eval/runs/hodge_cm_ablation.json Covers: Fase1 diag + offline ablation (CM/Hodge/RAND/HECKE) + functional probe.*

**Functions:**
- `resolve_device` (line 28) `def resolve_device(req)`
- `mat_stats` (line 48) `def mat_stats(kr, ki, r)`
- `heat_smooth` (line 64) `def heat_smooth(kr, ki, t, device)`
- `main` (line 74) `def main()`
- `sm` (line 70) `def sm(x)`
- `mean` (line 141) `def mean(k, m)`

#### `integration_smoke.py`
**Path:** `eval/integration_smoke.py`
**File Doc:** *End-to-end smoke test of all P0+P1 components working together.  Verifies that `run_one_test_sandboxed` can run a *valid* HumanEval candidate through the sandbox and get a pass=True result. This exercises the integration of: - sandbox.py (P0) - harness.py integration (the new run_one_test_sandboxed) - HumanEval canonical protocol (prompt + completion + test)*

**Functions:**
- `main` (line 18) `def main()`

#### `noise_analysis.py`
**Path:** `eval/noise_analysis.py`
**File Doc:** *Analisis post-hoc del noise sweep.  Generaciones del MISMO prompt bajo distintos niveles de ruido -> comparar con metricas que NO son pass@1 (porque los problemas triviales saturan):  - generation_exact_match:  % de generaciones que matchean exactamente el baseline (token por token) - prefix_agreement@50:    % de pares (baseline, noisy) que comparten el mismo prefijo de 50 tokens - levenshtein_dist:       distancia de edicion normalizada al baseline - token_jaccard:          interseccion / union de tokens generados - bleu_1:                 unigrama precision - syntax_ok_rate:         % que pasa ast.parse (sintaxis Python valida)  Salida: eval/runs/noise_analysis_<tag>.json*

**Functions:**
- `_load` (line 43) `def _load(p)`
- `consistency_across_runs` (line 47) `def consistency_across_runs(per_run)` - *Para cada problema, mira si pasa consistentemente a traves de los
4 niveles de ruido. Devuelve:
  - always_pass, always_fail, mixed (count)
  - per_sigma_pass_lists: {sigma: {task_id: bool}}*
- `main` (line 83) `def main()`

#### `noise_sweep.py`
**Path:** `eval/noise_sweep.py`
**File Doc:** *Barrido de ruido en los pesos espectrales del checkpoint TopoGPT3.  Para cada nivel sigma en --sigmas: 1. Carga el checkpoint base (NO modifica el archivo, solo los pesos en RAM) 2. Inyecta ruido gaussiano N(0, sigma) en los kernels espectrales cuaternionicos (kr_w/x/y/z, ki_w/x/y/z) y solo en ellos. Asi aislamos el efecto del ruido sobre la parte que el marco teorico dice que esta protegida topologicamente. 3. Genera pass@1 (greedy, T=0) sobre los primeros N problemas de HumanEval (subset para mantener tiempo de pared manejable) 4. Ejecuta los tests canonicos y mide pass rate  Salida: eval/runs/noise_<sigma>_<tag>.jsonl eval/runs/noise_sweep_<timestamp>.jsonl (resumen agregado)*

**Functions:**
- `inject_noise` (line 46) `def inject_noise(model, sigma, seed)` - *Anade N(0, sigma) a TODOS los kernels espectrales (kr_*, ki_*).
Retorna un dict con conteo de tensores ruidosos y de parametros
modificados.*
- `load_model` (line 74) `def load_model(ckpt_dir, ckpt_name, device)` - *Reconstruye TopoGPT2 alineado con el checkpoint, sin acceso a
harness.ModelLoader (queremos un loader limpio que no comparta
estado con corridas paralelas).*
- `generate_one` (line 99) `def generate_one(model, tok, prompt, max_new_tokens, device)`
- `main` (line 117) `def main()`

#### `repair.py`
**Path:** `eval/repair.py`
**File Doc:** *Self-repair loop on top of a greedy JSONL.  Takes the failed problems from --input, builds a rejection-feedback prompt that contains: - the original HumanEval prompt (signature + docstring) - the candidate the model wrote on its first attempt - the traceback from the hidden test - a "# fix:" cue  and re-prompts the model to rewrite the function. Runs N rounds. Each problem's *best* outcome across rounds is recorded.  Output: a new JSONL with the same shape as harness.py.*

**Functions:**
- `_new_loader` (line 36) `def _new_loader(ckpt_dir, ckpt_name)`
- `extract_candidate` (line 49) `def extract_candidate(prompt, completion)`
- `run_test` (line 75) `def run_test(problem, candidate_src)`
- `build_repair_prompt` (line 89) `def build_repair_prompt(prompt, candidate, err, entry_point)`
- `gen` (line 104) `def gen(model, tok, text, max_new_tokens, temperature, top_k, rep_penalty)`
- `main` (line 119) `def main()`

#### `report.py`
**Path:** `eval/report.py`
**File Doc:** *Aggregate every JSONL in eval/runs into a final report.  Reads runs from the original pass@k runs, the HRM run, the repair run, and produces: - pass@1 / pass@k tables - error-class breakdowns - wall-clock / throughput - a comparison standard vs HRM - a self-repair impact summary - emits REPORT.md next to the runs/*

**Functions:**
- `pass_at_k` (line 25) `def pass_at_k(n, c, k)`
- `classify_error` (line 31) `def classify_error(msg)`
- `load_jsonl` (line 52) `def load_jsonl(p)`
- `summarize_run` (line 56) `def summarize_run(p)`
- `repair_summary` (line 90) `def repair_summary(repair_path, baseline_path)`
- `main` (line 117) `def main()`

#### `samplers.py`
**Path:** `eval/samplers.py`
**File Doc:** *Registry of sampler constructors for the HumanEval harness.  Replaces the hardcoded `if mode == "standard": ... elif mode == "hrm": ...` chain in `eval.harness.make_sampler` with a decorator-based registry that mirrors the pattern in `claude-code-main/src/tools.ts`.  The pattern: - `@register_sampler("name")` decorates a factory function that takes a `settings_kwargs` dict (already filtered for sampler-specific keys) and returns an object with a `run()` method (or just a sampler object that `harness.evaluate_problem` knows how to drive). - `build_sampler("name", settings_kwargs)` is the public entry point. - `list_samplers()` returns the registered names for `--help` output.  Adding a new sampler is then a one-decorator change, not an edit to the harness's control flow.*

**Functions:**
- `register_sampler` (line 36) `def register_sampler(name)` - *Decorator. Register a factory under `name`. If `enabled_env` is set,
the factory is only registered when that env var is truthy. This
mirrors the `feature('XXX')` gating in claude-code-main/src/tools.ts.*
- `_is_env_truthy` (line 55) `def _is_env_truthy(name)`
- `_make_standard` (line 64) `def _make_standard(settings_kwargs)`
- `_make_hrm` (line 69) `def _make_hrm(settings_kwargs)`
- `list_samplers` (line 86) `def list_samplers()`
- `build_sampler` (line 90) `def build_sampler(mode, settings_kwargs)` - *Construct a sampler. Drop-in replacement for the old
`make_sampler(mode, settings_kwargs)` in `eval.harness`.*
- `deco` (line 42) `def deco(fn)`

#### `sandbox.py`
**Path:** `eval/sandbox.py`
**File Doc:** *Sandbox for executing model-generated code during HumanEval evaluation.  This module replaces the bare `exec()` call in `eval/harness.py:run_one_test` with a defence-in-depth check inspired by Claude Code's BashTool permission gates (see `claude-code-main/src/tools/BashTool/bashSecurity.ts`).  The threat model: - A language model emits Python source as a "candidate function". - The candidate is `exec()`'d alongside a hidden test. - Without protection, the model can `import os; os.system('rm -rf /')`, read secrets, fork-bomb, or hang the evaluator forever.  Layered defences (each can be disabled independently for debugging): 1. AST pre-check: parse the candidate, reject anything that imports dangerous modules, calls dangerous builtins, or shadows `__builtins__`. 2. Builtin whitelist: even if the candidate parses, `safe_exec` provides a stripped `__builtins__` without `open`, `exec`, `eval`, `__import__`, `compile`, `getattr` (controversial but standard). 3. Subprocess isolation: `safe_exec` runs the program in a child process so the OS enforces the timeout (vs. signal-based which the main thread can swallow). 4. Output capture: stdout/stderr are piped, not inherited from the parent terminal.  Usage: from eval.sandbox import safe_exec, check_safety, SandboxConfig  cfg = SandboxConfig(timeout=10.0, dry_run=False) ok, reason = check_safety(candidate_src, cfg) if not ok:*

**Classes:**
- `SandboxConfig` (line 53) `class SandboxConfig` - *One knob per defence layer. Defaults match HumanEval-style eval.*

**Methods:**
- `_names_imported` (line 100) `def _names_imported(tree)` - *Return the set of top-level names brought into scope by imports.*
- `_blocked_dunder_access` (line 114) `def _blocked_dunder_access(tree, blocked)` - *Find Attribute nodes whose attr is in `blocked`. Returns attr names found.*
- `_max_depth` (line 123) `def _max_depth(tree)` - *Compute max nesting depth of the AST. Catches obfuscated huge trees.*
- `check_safety` (line 133) `def check_safety(source, cfg)` - *Return (ok, reason). `reason` is "" when ok, else a human-readable
one-line explanation. Reasons are stable (used in test fixtures).*
- `_build_worker_src` (line 254) `def _build_worker_src(allowed_builtin_names, program_src, blocked_modules)`
- `safe_exec` (line 270) `def safe_exec(program_src, cfg, extra_globals)` - *Execute `program_src` in a sandboxed child process. Returns the same
5-tuple as `eval.harness.run_one_test` for drop-in compatibility.

The child is killed (SIGKILL) by the OS after `cfg.timeout` seconds.*
- `describe_policy` (line 373) `def describe_policy(cfg)`
- `d` (line 125) `def d(node, cur)`

#### `sandbox_smoke.py`
**Path:** `eval/sandbox_smoke.py`
**File Doc:** *Smoke test for eval.sandbox.  Verifies all four defence layers: L1 (AST pre-check): blocked imports & dunder attrs are rejected L2 (builtin whitelist): open/exec/etc raise NameError in the child L3 (subprocess isolation): infinite loops are killed at OS level L4 (output capture): stdout/stderr from the candidate are returned*

**Functions:**
- `main` (line 15) `def main()`

#### `smoke.py`
**Path:** `eval/smoke.py`
**File Doc:** *Smoke test: load the TopoGPT3 checkpoint and produce a small completion.  Used as the first gate: if this fails we abort HumanEval.*

**Functions:**
- `run_standard` (line 17) `def run_standard()`
- `run_hrm` (line 36) `def run_hrm()`

#### `temp_sweep.py`
**Path:** `eval/temp_sweep.py`
**File Doc:** *Barrido de temperatura x top-k sobre HumanEval.  Mide pass@1 en modo greedy (T=0) y pass@5 a temperaturas crecientes para mapear la "fase" de generacion del modelo:  - cristal: pass@1 alto, poca varianza entre samples - vidrio:  pass@1 bajo, alta varianza - caotico: pass@1 ~= 0, alta diversidad pero sin aciertos  Salida: eval/runs/temp_sweep_<tag>.jsonl (resumen) eval/runs/temp_<T>_top<k>_<tag>.jsonl (detalle por config)*

**Functions:**
- `generate_one` (line 39) `def generate_one(model, tok, prompt, max_new_tokens, temperature, top_k, device, seed_offset)`
- `evaluate_problems` (line 58) `def evaluate_problems(model, tok, problems, max_new_tokens, temperature, top_k, n_samples, device)`
- `pass_at_k_unbiased` (line 88) `def pass_at_k_unbiased(n, c, k)`
- `summarize` (line 96) `def summarize(results, n_samples)`
- `main` (line 116) `def main()`

#### `infer_exploitgym.py`
**Path:** `infer_exploitgym.py`
**File Doc:** *Standalone inference script for TopoExploit.  Loads the 125M TopoGPT-2 model and performs inference on exploit-related tasks using prompts derived from the .exploitgym_repo dataset.  Modes: --prompt "text"       Single prompt inference --eval-holdout        Evaluate on holdout tasks from the repo --interactive         Interactive prompt REPL  Tiers: vulnerability_analysis, patch_analysis, exploit_development*

**Functions:**
- `load_task_ids` (line 41) `def load_task_ids()`
- `load_task` (line 49) `def load_task(task_id)` - *Return a dict with keys: task_id, family, name, description, patch, pov.*
- `build_prompt` (line 122) `def build_prompt(task_info, tier)`
- `load_model` (line 136) `def load_model(ckpt_dir, ckpt_name, device)` - *Load tokenizer, build TopoGPT-2, apply Gauss patch, load weights.*
- `generate` (line 177) `def generate(model, tokenizer, prompt)` - *Autoregressive generation with streaming + n-gram repetition blocking.

Generates token by token using the model's KV cache. When ``stream`` is
True, decoded text is flushed incrementally (via ``stream_cb`` if given,
otherwise to stdout) so the user sees output as it is produced instead of
waiting for the whole sequence. Repeated n-grams of size
``no_repeat_ngram`` are suppressed to avoid degenerate repetition loops.*
- `run_single_prompt` (line 257) `def run_single_prompt(args, model, tokenizer)`
- `run_eval_holdout` (line 274) `def run_eval_holdout(args, model, tokenizer)`
- `run_interactive` (line 331) `def run_interactive(args, model, tokenizer)`
- `parse_args` (line 380) `def parse_args()`
- `main` (line 420) `def main()`
- `sample` (line 205) `def sample(logits, seen_ids)`

#### `synthetic_dataset.py`
**Path:** `synthetic_dataset.py`
**File Doc:** *Synthetic Dataset Generator for TopoGPT2.  Generates high-quality code instruction-tuning data from existing source files using a multi-stage LLM pipeline:  file → analysis → spec → chain-of-thought → vague question → JSONL  Each sample (JSONL line) contains: { "instruction":  "vague natural question", "thinking":     "chain-of-thought reasoning", "spec":         "detailed spec-driven prompt", "todo":         ["task 1", "task 2", ...], "response":     "```language\noriginal clean code\n```", "file_path":    "src/foo/bar.py", "lang":         "python", "checksum":     "sha256 of original code", }  Pipeline is designed for efficiency: - One LLM call per file (master prompt, one-shot) - Streaming JSONL writes (never holds full dataset in memory) - SHA256 dedup across the full corpus - Resumable: tracks processed files in a manifest - Batch-friendly: process N files per run  Backend: Groq API (Llama-3.3-70B, fastest/cheapest) or OpenRouter.*

**Classes:**
- `LLMBackend` (line 61) `class LLMBackend` - *Abstract LLM backend. Subclass for each provider.*
- `GroqBackend` (line 71) `class GroqBackend(LLMBackend)` - *Groq API backend using requests.

Supports models: llama-3.3-70b-versatile, deepseek-r1.
Set GROQ_API_KEY env var.*
- `OpenRouterBackend` (line 121) `class OpenRouterBackend(LLMBackend)` - *OpenRouter unified API backend.

Supports any OpenRouter model:
    anthropic/claude-3.5-sonnet,
    openai/gpt-4o,
    deepseek/deepseek-chat,
    google/gemini-2.0-flash-thinking,
Set OPENROUTER_API_KEY env var.*
- `OllamaBackend` (line 177) `class OllamaBackend(LLMBackend)` - *Ollama local inference backend.

Supports any local model: llama3.1:8b, granite4.1:3b, etc.
Connects to Ollama server at OLLAMA_HOST (default: http://localhost:11434).*
- `ProcessedManifest` (line 364) `class ProcessedManifest` - *Tracks processed files for resumability.*
- `SyntheticDatasetGenerator` (line 399) `class SyntheticDatasetGenerator` - *Generates synthetic instruction-tuning data from source files.

Pipeline (one LLM call per file):
    file → MASTER_PROMPT → LLM → validate → dedup → JSONL

Features:
- Streaming JSONL writes (bounded RAM)
- SHA256 dedup across corpus
- Resumable (manifest tracks progress)
- Threaded request batching for throughput
- Configurable quality thresholds*

**Methods:**
- `build_backend` (line 227) `def build_backend(provider, model)` - *Factory for LLM backends.*
- `validate_sample` (line 330) `def validate_sample(sample)` - *Validate that a generated sample meets quality bar.

Returns (is_valid, reason).*
- `build_logger` (line 614) `def build_logger(level)`
- `parse_args` (line 625) `def parse_args()`
- `load_paths` (line 652) `def load_paths(paths_arg, paths_file, max_files)` - *Load file paths from CLI args or file.*
- `main` (line 667) `def main()`
- `generate` (line 64) `def generate(self, prompt)`
- `name` (line 67) `def name(self)`
- `__init__` (line 78) `def __init__(self, model, api_key, max_tokens, temperature, timeout)`
- `name` (line 95) `def name(self)`
- `generate` (line 98) `def generate(self, prompt)`
- `__init__` (line 132) `def __init__(self, model, api_key, max_tokens, temperature, timeout)`
- `name` (line 151) `def name(self)`
- `generate` (line 154) `def generate(self, prompt)`
- `__init__` (line 184) `def __init__(self, model, host, max_tokens, temperature, timeout)`
- `name` (line 198) `def name(self)`
- `generate` (line 201) `def generate(self, prompt)`
- `load` (line 374) `def load(path)`
- `save` (line 387) `def save(self, path)`
- `__init__` (line 418) `def __init__(self, backend, output_path, manifest_path, logger, max_workers, max_file_chars)`
- `_jsonl_writer` (line 447) `def _jsonl_writer(self)` - *Background thread that drains the queue and writes JSONL lines.*
- `_enqueue_sample` (line 465) `def _enqueue_sample(self, sample)`
- `_flush_writer` (line 468) `def _flush_writer(self)`
- `_read_file` (line 477) `def _read_file(self, path)` - *Read file content and detect language. Truncate if needed.*
- `_build_prompt` (line 490) `def _build_prompt(self, content, lang)`
- `_generate_sample` (line 496) `def _generate_sample(self, content, lang)` - *Call LLM with retry logic.*
- `process_file` (line 533) `def process_file(self, path)` - *Process a single file. Returns True if a sample was written.*
- `process_batch` (line 568) `def process_batch(self, paths)` - *Process a batch of files in parallel using thread pool.*
- `finish` (line 590) `def finish(self)` - *Signal end of processing and flush writer.*

#### `test_jlens.py`
**Path:** `tests/test_jlens.py`

**Classes:**
- `TestValidPositionMask` (line 17) `class TestValidPositionMask` - *Feature: valid_position_mask excludes attention-sink and final positions.*
- `TestJacobianForPrompt` (line 52) `class TestJacobianForPrompt` - *Feature: jacobian_for_prompt computes J_l for one prompt.*
- `TestFit` (line 172) `class TestFit` - *Feature: fit() averages Jacobians over multiple prompts.*
- `TestJacobianLens` (line 210) `class TestJacobianLens` - *Feature: JacobianLens saves, loads, applies, and merges.*
- `TestFitCheckpoint` (line 367) `class TestFitCheckpoint` - *Feature: fit() with checkpoint resume works correctly.*
- `TestConfig` (line 473) `class TestConfig` - *Feature: Config classes centralize all tunable parameters.*
- `TestTopoGPT3JLensAppConfig` (line 494) `class TestTopoGPT3JLensAppConfig` - *Feature: Application config controls readout behavior.*

**Methods:**
- `test_basic_mask` (line 20) `def test_basic_mask(self)` - *Scenario: Correct mask for a standard-length prompt.*
- `test_too_short_raises` (line 29) `def test_too_short_raises(self)` - *Scenario: Too-short prompt raises ValueError.*
- `test_negative_skip_raises` (line 34) `def test_negative_skip_raises(self)` - *Scenario: Negative skip_first raises ValueError.*
- `test_all_positions_valid` (line 39) `def test_all_positions_valid(self)` - *Scenario: skip_first=0 includes all but final position.*
- `test_exact_minimum_length` (line 45) `def test_exact_minimum_length(self)` - *Scenario: Exact minimum length (skip_first + 2) works.*
- `model` (line 56) `def model(self)`
- `test_returns_jacobians_for_source_layers` (line 63) `def test_returns_jacobians_for_source_layers(self, model)` - *Scenario: Returns Jacobians for all requested source layers.*
- `test_late_layer_jacobian_close_to_identity` (line 76) `def test_late_layer_jacobian_close_to_identity(self, model)` - *Scenario: J_{n_layers-2} has diag ~= 1 (identity property).*
- `test_earlier_layers_further_from_identity` (line 85) `def test_earlier_layers_further_from_identity(self, model)` - *Scenario: Earlier layers compound deviations from identity.*
- `test_exact_jacobian_for_last_block` (line 95) `def test_exact_jacobian_for_last_block(self, model)` - *Scenario: J_{n_layers-2} equals I + W_{last} exactly.

For TinyDecoder with block = h + 0.1*W*h, J_{n_layers-2} = I + W.*
- `test_negative_layer_indices` (line 110) `def test_negative_layer_indices(self, model)` - *Scenario: Negative layer indices are normalized correctly.*
- `test_out_of_range_layers_rejected` (line 133) `def test_out_of_range_layers_rejected(self, model)` - *Scenario: Out-of-range layers raise ValueError.*
- `test_source_below_target_enforced` (line 145) `def test_source_below_target_enforced(self, model)` - *Scenario: source_layers must be below target_layer.*
- `test_target_out_of_range_raises` (line 158) `def test_target_out_of_range_raises(self, model)` - *Scenario: target_layer out of range raises ValueError.*
- `model` (line 176) `def model(self)`
- `test_fit_returns_lens_with_correct_attributes` (line 183) `def test_fit_returns_lens_with_correct_attributes(self, model)` - *Scenario: fit() returns JacobianLens with correct metadata.*
- `test_fit_empty_prompts_raises` (line 191) `def test_fit_empty_prompts_raises(self, model)` - *Scenario: No valid prompts raises ValueError.*
- `test_fit_skips_short_prompts` (line 196) `def test_fit_skips_short_prompts(self, model)` - *Scenario: Too-short prompts are skipped.*
- `test_fit_with_default_source_layers` (line 202) `def test_fit_with_default_source_layers(self, model)` - *Scenario: Default source_layers covers all layers below target.*
- `model` (line 214) `def model(self)`
- `fitted_lens` (line 222) `def fitted_lens(self, model)`
- `test_save_and_load_round_trip` (line 226) `def test_save_and_load_round_trip(self, fitted_lens, tmp_path)` - *Scenario: save/load preserves jacobians (fp16 tolerance).*
- `test_apply_returns_correct_shapes` (line 242) `def test_apply_returns_correct_shapes(self, fitted_lens, model)` - *Scenario: apply() returns correct logit shapes.*
- `test_fitted_late_layer_matches_model` (line 254) `def test_fitted_late_layer_matches_model(self, fitted_lens, model)` - *Scenario: Transported late-layer logits match model logits.*
- `test_apply_with_explicit_positions` (line 263) `def test_apply_with_explicit_positions(self, fitted_lens, model)` - *Scenario: Explicit positions return correct subset.*
- `test_logit_lens_baseline` (line 274) `def test_logit_lens_baseline(self, fitted_lens, model)` - *Scenario: use_jacobian=False returns untransported logits.*
- `test_unfitted_layer_rejected` (line 281) `def test_unfitted_layer_rejected(self, fitted_lens, model)` - *Scenario: Unfitted layer raises ValueError.*
- `test_out_of_range_layer_rejected` (line 286) `def test_out_of_range_layer_rejected(self, fitted_lens, model)` - *Scenario: Out-of-range layer raises ValueError.*
- `test_merge_weighted_mean` (line 291) `def test_merge_weighted_mean(self)` - *Scenario: merge() computes n_prompts-weighted mean.*
- `test_merge_mismatch_raises` (line 319) `def test_merge_mismatch_raises(self)` - *Scenario: Mismatched lenses raise ValueError.*
- `test_merge_empty_raises` (line 326) `def test_merge_empty_raises(self)` - *Scenario: Empty merge raises ValueError.*
- `test_transport_produces_correct_shape` (line 331) `def test_transport_produces_correct_shape(self, fitted_lens)` - *Scenario: transport() maps residual to final-layer basis.*
- `test_load_invalid_file_raises` (line 337) `def test_load_invalid_file_raises(self, tmp_path)` - *Scenario: Loading non-lens file raises ValueError.*
- `test_from_pretrained_local_file` (line 344) `def test_from_pretrained_local_file(self, fitted_lens, tmp_path)` - *Scenario: from_pretrained resolves a local file.*
- `test_from_pretrained_local_directory` (line 351) `def test_from_pretrained_local_directory(self, fitted_lens, tmp_path)` - *Scenario: from_pretrained resolves a local directory.*
- `test_repr` (line 359) `def test_repr(self, fitted_lens)` - *Scenario: repr contains key metadata.*
- `model` (line 371) `def model(self)`
- `test_checkpoint_resume_produces_same_result` (line 378) `def test_checkpoint_resume_produces_same_result(self, model, tmp_path)` - *Scenario: Resumed fit matches fresh fit.*
- `test_resume_after_skip_no_double_count` (line 408) `def test_resume_after_skip_no_double_count(self, model, tmp_path)` - *Scenario: Resume after a skipped prompt does not double-count.

Regression: a skipped prompt must not desync success-count from
list-position.*
- `test_checkpoint_mismatch_raises` (line 450) `def test_checkpoint_mismatch_raises(self, model, tmp_path)` - *Scenario: Mismatched checkpoint settings raise ValueError.*
- `test_fit_config_defaults` (line 476) `def test_fit_config_defaults(self)` - *Scenario: Default fit config has sensible defaults.*
- `test_app_config_defaults` (line 485) `def test_app_config_defaults(self)` - *Scenario: Default app config has sensible defaults.*
- `test_default_config` (line 497) `def test_default_config(self)` - *Scenario: Default app config uses all positions.*
- `test_custom_config` (line 505) `def test_custom_config(self)` - *Scenario: Custom app config overrides specific layers.*

#### `test_lens_model.py`
**Path:** `tests/test_lens_model.py`

**Classes:**
- `TestTopoGPT3LensConfig` (line 13) `class TestTopoGPT3LensConfig` - *Feature: TopoGPT3LensConfig provides centralized adapter configuration.*
- `TestTinyDecoder` (line 41) `class TestTinyDecoder` - *Feature: TinyDecoder provides a minimal test model.*
- `TestTopoGPT3LensModel` (line 65) `class TestTopoGPT3LensModel` - *Feature: TopoGPT3LensModel wraps a model to implement LensModel protocol.*
- `TestTopoGPT3LensModelWithRecording` (line 214) `class TestTopoGPT3LensModelWithRecording` - *Feature: ActivationRecorder works with TopoGPT3LensModel.*
- `TestTopoGPT3LensModelEdgeCases` (line 278) `class TestTopoGPT3LensModelEdgeCases` - *Feature: Edge cases are handled gracefully.*

**Methods:**
- `test_default_config` (line 16) `def test_default_config(self)` - *Scenario: Default config matches small scale preset.*
- `test_from_topogpt2_config` (line 25) `def test_from_topogpt2_config(self)` - *Scenario: Build lens config from TopoGPT2Config.*
- `test_probe_checkpoint_missing_raises` (line 35) `def test_probe_checkpoint_missing_raises(self, tmp_path)` - *Scenario: Missing state.json raises FileNotFoundError.*
- `test_default_parameters` (line 44) `def test_default_parameters(self)` - *Scenario: TinyDecoder has correct default shape.*
- `test_forward_output_shape` (line 51) `def test_forward_output_shape(self)` - *Scenario: Forward pass produces correct logit shape.*
- `test_weight_tied` (line 59) `def test_weight_tied(self)` - *Scenario: Embedding and LM head share weights.*
- `raw_model` (line 69) `def raw_model(self)`
- `lens_model` (line 77) `def lens_model(self, raw_model)`
- `test_exposes_protocol_attributes` (line 80) `def test_exposes_protocol_attributes(self, lens_model, raw_model)` - *Scenario: LensModel attributes match underlying model.*
- `test_encode_text_to_token_ids` (line 87) `def test_encode_text_to_token_ids(self, lens_model)` - *Scenario: encode() returns tensor of shape [1, seq_len].*
- `test_encode_with_tokenizer` (line 95) `def test_encode_with_tokenizer(self)` - *Scenario: encode() uses BPETokenizer when available.*
- `test_encode_respects_max_length` (line 107) `def test_encode_respects_max_length(self, lens_model)` - *Scenario: encode() truncates at max_length.*
- `test_forward_returns_residual_only` (line 113) `def test_forward_returns_residual_only(self)` - *Scenario: forward() returns hidden states with d_model dim, not vocab.

The lens model forward should stop before final_norm and lm_head.
The output should have d_model as last dimension, not vocab_size.*
- `test_forward_differs_from_full_model` (line 128) `def test_forward_differs_from_full_model(self)` - *Scenario: Residual forward shape differs from full model logits.*
- `test_unembed_produces_logits` (line 141) `def test_unembed_produces_logits(self, lens_model)` - *Scenario: unembed() maps residual to logits.*
- `test_forward_plus_unembed_matches_model_logits` (line 150) `def test_forward_plus_unembed_matches_model_logits(self, lens_model, raw_model)` - *Scenario: residual forward + unembed == model forward logits.

This validates that our split forward matches the original model's
full forward pass.*
- `test_autograd_graph_tracks_through_layers` (line 163) `def test_autograd_graph_tracks_through_layers(self)` - *Scenario: Gradient flows through residual layers when grads enabled.*
- `test_input_device_property` (line 180) `def test_input_device_property(self, lens_model)` - *Scenario: input_device returns the embedding weight device.*
- `test_input_device_setter` (line 185) `def test_input_device_setter(self, lens_model)` - *Scenario: input_device can be overridden.*
- `test_tokenizer_setter` (line 191) `def test_tokenizer_setter(self, lens_model)` - *Scenario: tokenizer can be set after construction.*
- `test_from_checkpoint_missing_raises` (line 198) `def test_from_checkpoint_missing_raises(self)` - *Scenario: from_checkpoint with missing directory raises.*
- `test_grad_enabled_deterministic` (line 205) `def test_grad_enabled_deterministic(self, lens_model)` - *Scenario: Multiple forward passes with same input are deterministic.*
- `lens_model` (line 218) `def lens_model(self)`
- `test_recorder_captures_layer_outputs` (line 225) `def test_recorder_captures_layer_outputs(self, lens_model)` - *Scenario: ActivationRecorder captures all requested layer outputs.*
- `test_recorder_with_start_graph_at` (line 238) `def test_recorder_with_start_graph_at(self, lens_model)` - *Scenario: start_graph_at roots the autograd graph.*
- `test_recorder_cleanup_on_exception` (line 252) `def test_recorder_cleanup_on_exception(self, lens_model)` - *Scenario: Hooks are removed even if construction fails.*
- `test_recorder_detach_after_forward` (line 264) `def test_recorder_detach_after_forward(self, lens_model)` - *Scenario: Activations can be detached after recorder exits.*
- `test_empty_sequence` (line 281) `def test_empty_sequence(self)` - *Scenario: Empty input produces error or minimal output.*
- `test_single_token` (line 291) `def test_single_token(self)` - *Scenario: Single token input works.*

#### `__init__.py`
**Path:** `topogpt3/__init__.py`
**File Doc:** *TopoGPT3: complex-valued spectral language model for code.  This package bundles:  - ``topogpt3.model``: the base TopoGPT2 architecture (quaternion spectral layers, BPE tokenizer, helpers). - ``topogpt3.train``: the curriculum trainer with Grassmannian / Fisher / phase diagnostics. - ``topogpt3.inference``: a standard autoregressive sampler that loads a trained safetensors checkpoint. - ``topogpt3.inference_hrm``: a hierarchical recursive reasoning sampler that reuses the same checkpoint with no extra trained parameters. - ``topogpt3.lens_model``: the Jacobian-lens model adapter (LensModel protocol + TopoGPT3LensModel wrapper). - ``topogpt3.jlens``: Jacobian lens fitting, application, and the ActivationRecorder / JacobianLens infrastructure.  Typical usage from a downstream project::  from topogpt3 import InferenceSettings, InferencePipeline  settings = InferenceSettings( checkpoint_dir="checkpoints_topogpt3", prompt="def fibonacci(", max_new_tokens=200, ) InferencePipeline(settings).execute()  Jacobian lens usage::*

*No symbols extracted*

#### `__main__.py`
**Path:** `topogpt3/__main__.py`

**Functions:**
- `main` (line 6) `def main()` - *TopoGPT3 entry point. Delegates to subcommands.*

#### `api_server.py`
**Path:** `topogpt3/api_server.py`
**File Doc:** *OpenAI-compatible HTTP API server so TopoGPT3 can be used as a backend for coding agents (e.g. Pi, Aider, Continue, Codex CLI, etc.).  Security Posture ---------------- - **Authentication**: Bearer token (``Authorization: Bearer <key>``). Keys are loaded from ``--keys`` (comma-separated) or the ``TOPOGPT3_API_KEYS`` env var. Admin keys (prefixed ``admin:``) get higher rate limits. Constant-time comparison prevents timing leaks. - **Authorization**: token-bucket rate limiter per-key and per-IP with configurable thresholds. After ``max_failures`` bad auth attempts an IP is banned for ``ban_window`` seconds. - **Input hardening**: Pydantic schemas enforce strict types, min/max bounds, and length limits. Request body is capped server-side. Error responses never leak stack traces. - **Headers**: ``X-Content-Type-Options: nosniff``, ``X-Frame-Options: DENY``, ``X-XSS-Protection: 1; mode=block``, ``Content-Security-Policy: default-src 'none'`` on every response. CORS policy allows nothing by default (configurable allow-origins). - **Audit**: structured JSON log lines for every request (truncated bodies, no secrets).  Usage::  TOPOGPT3_API_KEYS="sk-secret-key,admin:sk-admin-key" \\ python -m topogpt3 api_server \\ --checkpoint checkpoints_topogpt3/last \\ --port 8800  Pi / agent config::*

**Classes:**
- `ApiKey` (line 137) `class ApiKey`
- `AuthState` (line 143) `class AuthState`
- `TokenBucket` (line 202) `class TokenBucket`
- `RateLimiter` (line 219) `class RateLimiter`
- `IpBanner` (line 250) `class IpBanner`
- `CompletionRequest` (line 291) `class CompletionRequest(BaseModel)`
- `Message` (line 310) `class Message(BaseModel)`
- `ChatCompletionRequest` (line 316) `class ChatCompletionRequest(BaseModel)`
- `ServerModel` (line 341) `class ServerModel`

**Functions:**
- `_setup_logging` (line 116) `def _setup_logging(verbose)`

**Methods:**
- `_parse_keys` (line 164) `def _parse_keys(raw)` - *Accept ``key1,admin:key2,key3``. The ``admin:`` prefix marks an
admin-level key; everything else is a regular user key.*
- `_sha256` (line 192) `def _sha256(raw)`
- `_sanitize_stop` (line 281) `def _sanitize_stop(stop)`
- `_resolve_device` (line 499) `def _resolve_device(device)`
- `_probe_arch` (line 505) `def _probe_arch(checkpoint_dir)` - *Probe a checkpoint and return (scale, n_kv_heads).

Reads the model hidden dim from ``final_norm.weight`` (or the first
transformer layer's output projection) and the KV head count from the
``k_proj`` output dim, so the server can load checkpoints of any scale
(small/large/…) without hardcoding the preset.*
- `load_model` (line 553) `def load_model(checkpoint, device)`
- `lifespan` (line 577) `def lifespan(app)`
- `_security_middleware` (line 617) `def _security_middleware(request, call_next)` - *Global middleware: rate-limit, IP-ban, security headers, audit log.*
- `_real_ip` (line 645) `def _real_ip(request)` - *Best-effort real client IP. We trust no proxy headers by default.*
- `_json_error` (line 656) `def _json_error(status, detail)`
- `_authenticate` (line 668) `def _authenticate(request)` - *FastAPI dependency: extract & validate Bearer token.

When no API keys are configured (auth disabled via ``--no-auth`` or an
empty ``TOPOGPT3_API_KEYS``), the dependency is a pass-through so local
integrations (e.g. the ExploitGym TopoExploit agent) can call it without
a token.*
- `_check_rate_limit` (line 694) `def _check_rate_limit(api_key, request)` - *Rate limit per-key (with admin exemption / higher limit).*
- `health` (line 712) `def health(request)`
- `list_models` (line 719) `def list_models(request)`
- `completions` (line 736) `def completions(req, request)`
- `chat_completions` (line 792) `def chat_completions(req, request)`
- `_check_model` (line 850) `def _check_model()`
- `_short_id` (line 855) `def _short_id()`
- `_build_chat_prompt` (line 859) `def _build_chat_prompt(messages)`
- `_extract_text` (line 866) `def _extract_text(content)`
- `_stream_completion` (line 880) `def _stream_completion(prompt, max_tokens, temperature, top_k, repetition_penalty, stop, auto_continue, max_continuations)`
- `_stream_chat` (line 915) `def _stream_chat(t0_ms, prompt, max_tokens, temperature, top_k, repetition_penalty, stop, auto_continue, max_continuations)`
- `main` (line 953) `def main()`
- `validate` (line 148) `def validate(self, raw)`
- `consume` (line 208) `def consume(self, n)`
- `__init__` (line 220) `def __init__(self, user_rps, admin_rps, capacity)`
- `_cleanup` (line 227) `def _cleanup(self)`
- `allow` (line 233) `def allow(self, key, role)`
- `__init__` (line 251) `def __init__(self, max_failures, window)`
- `record_failure` (line 257) `def record_failure(self, ip)`
- `is_banned` (line 265) `def is_banned(self, ip)`
- `_normalize_stop` (line 306) `def _normalize_stop(cls, v)`
- `_normalize_stop` (line 331) `def _normalize_stop(cls, v)`
- `complete` (line 348) `def complete(self, prompt)`
- `stream_complete` (line 393) `def stream_complete(self, prompt)`
- `_is_eos` (line 482) `def _is_eos(self, token_id)`

#### `continuation.py`
**Path:** `topogpt3/continuation.py`
**File Doc:** *Auto-continuation engine: detects truncated responses and feeds the last incomplete lines back so the model can resume where it left off.  Used by both the standard inference pipeline and the HRM "thinking" mode.*

**Functions:**
- `_count_unclosed_brackets` (line 25) `def _count_unclosed_brackets(text)`
- `_count_unclosed_fences` (line 36) `def _count_unclosed_fences(text)`
- `is_response_complete` (line 45) `def is_response_complete(text, min_chars)` - *Heuristic to decide whether a model response looks finished.

Returns True when the response seems naturally complete (no need to
continue), False when it appears truncated and continuation may help.*
- `extract_tail_for_continuation` (line 75) `def extract_tail_for_continuation(text, tail_lines, tail_chars)` - *Return the last N lines (or up to tail_chars) of `text` as a
continuation prefix to feed back into the model.

The returned string can be prepended as context for the model's next
generation call so it continues naturally from that point.*
- `split_at_last_newline` (line 105) `def split_at_last_newline(text)` - *Split `text` at the last newline.

Returns (prefix_without_last_line, last_line).
Useful for discarding a trailing incomplete line before continuation.*

#### `ewc.py`
**Path:** `topogpt3/ewc.py`
**File Doc:** *Elastic Weight Consolidation (EWC) and Experience Replay for TopoGPT3.  Provides EWCRegularizer for regularizing against catastrophic forgetting across curriculum tiers, and ReplayBuffer for interleaving data from previous tiers during training.  Author: Gris Iscomeback License: GPL v3*

**Classes:**
- `EWCRegularizer` (line 25) `class EWCRegularizer` - *Elastic Weight Consolidation regularizer.

Computes the diagonal Fisher information matrix over a dataset and uses
it to penalize large deviations from the parameters that were optimal
for previous tasks (tiers). The penalty is:

    L_EWC = lambda * sum_i  F_i * (theta_i - theta_opt_i)^2

where F_i is the Fisher information for parameter i and theta_opt_i
is the parameter value at the end of the previous task.

Usage:
    ewc = EWCRegularizer(model, lambda_ewc=5000.0)
    ewc.compute_fisher(dataloader, vocab_size, device)
    # ... train on new tier ...
    loss = task_loss + ewc.penalty(model)*
- `ReplayBuffer` (line 223) `class ReplayBuffer` - *Experience Replay buffer for interleaving data from previous curriculum
tiers during training.

Maintains references to the training dataloaders of previously completed
tiers and mixes their batches into the current tier's training loop at a
configurable ratio.

Usage:
    replay = ReplayBuffer(max_tiers=3, replay_ratio=0.25)
    replay.add_tier(0, tier0_train_dl)
    replay.add_tier(1, tier1_train_dl)
    # During tier-2 training, for each step:
    bx, by = replay.sample_batch(current_batch_size=4)
    # bx, by contains replay data from tiers 0,1*

**Methods:**
- `__init__` (line 45) `def __init__(self, model, lambda_ewc, fisher_num_samples)` - *Args:
    model: The model to regularize.
    lambda_ewc: Penalty strength scaling the Fisher-weighted
        parameter deviation loss.
    fisher_num_samples: Number of batches used when computing the
        Fisher information matrix (stored for reference).*
- `compute_fisher` (line 68) `def compute_fisher(self, dataloader, vocab_size, device, max_batches)` - *Compute the diagonal Fisher information matrix and snapshot the
current parameters as the reference point for the EWC penalty.

For each batch: forward pass -> per-sample NLL -> backprop ->
accumulate squared gradients.  Handles AMP autocast.

After calling this method, ``self.fisher`` and ``self.old_params``
are populated and ready for ``penalty()``.

Args:
    dataloader: Iterable yielding ``(bx, by)`` batches.
    vocab_size: Size of the vocabulary (for reshaping logits).
    device: Device where the model lives.
    max_batches: Maximum number of batches to use for estimation.*
- `penalty` (line 168) `def penalty(self, model)` - *Compute the EWC regularization loss.

Returns:
    Scalar tensor: lambda * sum_i F_i * (theta_i - theta_opt_i)^2*
- `save_state` (line 186) `def save_state(self, path)` - *Persist the Fisher information matrix and old parameter snapshots
to a single ``.pt`` file.

Args:
    path: Destination file path (should end in ``.pt``).*
- `load_state` (line 203) `def load_state(self, path)` - *Load a previously saved EWC state.

Args:
    path: Path to the ``.pt`` file created by ``save_state``.*
- `__init__` (line 241) `def __init__(self, max_tiers, replay_ratio)` - *Args:
    max_tiers: Maximum number of previous tiers to remember.
        When exceeded, the oldest tier is dropped.
    replay_ratio: Fraction of the returned batch that comes from
        replayed tiers (0.0 = no replay, 1.0 = all replay).*
- `add_tier` (line 259) `def add_tier(self, tier_index, dataloader)` - *Register a tier's training dataloader for future replay.

If the buffer is full (``max_tiers`` tiers stored), the oldest
entry is evicted.

Args:
    tier_index: Integer tier identifier (e.g. 0, 1, 2, ...).
    dataloader: The training dataloader for that tier.*
- `_get_next_batch` (line 282) `def _get_next_batch(self, tier_index, dataloader)` - *Get the next batch from a tier's dataloader, cycling if necessary.

Args:
    tier_index: The tier identifier.
    dataloader: The tier's dataloader.

Returns:
    ``(bx, by)`` tensors or ``None`` if the dataloader is empty.*
- `sample_batch` (line 312) `def sample_batch(self, current_batch_size)` - *Sample a mixed replay batch from ALL stored previous tiers.

The returned batch has size approximately
``replay_ratio * current_batch_size`` and is drawn uniformly from
all stored tiers, cycling through their dataloaders.

The caller is responsible for concatenating this with the current
tier's batch::

    current_batch = next(current_dataloader)
    replay = replay_buffer.sample_batch(current_batch_size)
    if replay is not None:
        bx = torch.cat([current_batch[0], replay[0]], dim=0)
        by = torch.cat([current_batch[1], replay[1]], dim=0)

Args:
    current_batch_size: The batch size of the current tier.
        Used to compute the target replay batch size as
        ``max(1, int(replay_ratio * current_batch_size))``.

Returns:
    ``(bx, by)`` tensors on the same device as the dataloader
    yields them, or ``None`` if no replay tiers are available.*

#### `exploitgym_config.py`
**Path:** `topogpt3/exploitgym_config.py`
**File Doc:** *TopoExploit Training Configuration  122M parameter model trained on ExploitGym vulnerability/exploit data.  Tier structure: Tier 1: vulnerability_analysis (256 seq, 3 epochs) Tier 2: patch_analysis        (512 seq, 4 epochs) Tier 3: exploit_development   (1024 seq, 5 epochs)  Author: Gris Iscomeback License: GPL v3*

**Classes:**
- `TopoExploitConfig` (line 26) `class TopoExploitConfig` - *Configuration for TopoExploit: 122M params on ExploitGym data.*

**Methods:**
- `build_topogpt2_config` (line 100) `def build_topogpt2_config(self, max_seq_len, attn_window)`
- `build_loader_config` (line 122) `def build_loader_config(self)`

#### `exploitgym_loader.py`
**Path:** `topogpt3/exploitgym_loader.py`
**File Doc:** *ExploitGym Data Loader for TopoExploit  Clones the ExploitGym repo and tokenizes task data into training memmaps. Each task contains vulnerability descriptions, patches, exploit docs, and PoC source.  Tier structure: Tier 1: vulnerability_analysis -- CVE description -> vulnerability analysis Tier 2: patch_analysis        -- source + patch -> fix explanation Tier 3: exploit_development   -- vulnerability -> exploit strategy + PoC code  Author: Gris Iscomeback License: GPL v3*

**Classes:**
- `ExploitGymLoaderConfig` (line 39) `class ExploitGymLoaderConfig`
- `ExploitGymDataLoader` (line 200) `class ExploitGymDataLoader`

**Methods:**
- `ensure_repo` (line 54) `def ensure_repo(repo_cache, logger)`
- `_read_file_safe` (line 75) `def _read_file_safe(path, max_chars)`
- `_parse_task_ids` (line 83) `def _parse_task_ids(repo_path)`
- `_task_id_to_path` (line 93) `def _task_id_to_path(task_id)`
- `parse_task` (line 101) `def parse_task(task_dir, task_id, max_chars)`
- `_format_vulnerability_analysis` (line 137) `def _format_vulnerability_analysis(task)`
- `_format_patch_analysis` (line 153) `def _format_patch_analysis(task)`
- `_format_exploit_development` (line 168) `def _format_exploit_development(task)`
- `__init__` (line 201) `def __init__(self, config, tokenizer, logger)`
- `_tier_paths` (line 209) `def _tier_paths(self, tier)`
- `_manifest_path` (line 215) `def _manifest_path(self, tier)`
- `_already_prepared` (line 218) `def _already_prepared(self, tier)`
- `_load_all_tasks` (line 240) `def _load_all_tasks(self, repo_path)`
- `prepare_tier` (line 264) `def prepare_tier(self, tier_index, force)`
- `open_memmap` (line 411) `def open_memmap(self, tier, split)`
- `stratum_key` (line 288) `def stratum_key(task)`
- `flush` (line 343) `def flush(split)`

#### `hodge_cm.py`
**Path:** `topogpt3/hodge_cm.py`
**File Doc:** *Fases 2-5: martillo Hodge-CM para TopoGPT3 / modelo espectral continuo. F2: CM phase-lattice STE suave. F3: Hodge regularizer + Fisher hinge. F4: MCMC offline. F5: Hecke pooling solo escala.*

**Classes:**
- `CMPhaseQuantSTE` (line 13) `class CMPhaseQuantSTE(Function)`
- `HeckeScalePool` (line 147) `class HeckeScalePool(Module)` - *Promedio sobre orbitas de escala/traslacion. Solo donde hay simetria real
(Strassen T, Fourier interp Willmore). NO reemplazo de atencion general.*

**Methods:**
- `cm_phase_loss` (line 25) `def cm_phase_loss(kr, ki, m)` - *E_Q: distancia cuadratica media a raices m-esimas. kr,ki reales misma forma.*
- `apply_cm_soft` (line 34) `def apply_cm_soft(kr, ki, m, beta)` - *Cuantizacion suave de fase con STE. Magnitud intacta.*
- `freq_grid_laplacian` (line 49) `def freq_grid_laplacian(fh, fw, device, dtype)`
- `hodge_heat_residue` (line 66) `def hodge_heat_residue(kr, ki, t)` - *||K - e^{-tL} K||_F^2 sobre grid freq. kr,ki: [...,Fh,Fw].*
- `harmonic_ratio` (line 81) `def harmonic_ratio(kr, ki, t)`
- `fisher_hinge` (line 96) `def fisher_hinge(singulars, r, margin)`
- `langevin_refine` (line 105) `def langevin_refine(param, energy_fn, steps, eta, temp, seed)` - *ULA offline: theta <- theta - eta grad E + sqrt(2 eta T) xi.*
- `gibbs_cluster_assign` (line 123) `def gibbs_cluster_assign(feats, n_clusters, iters, seed)` - *Clustering blando por fase para nodos/frecuencias. feats: [N,D] complejo o real.*
- `forward` (line 15) `def forward(ctx, phase, m, beta)`
- `backward` (line 21) `def backward(ctx, grad_out)`
- `idx` (line 53) `def idx(h, w)`
- `__init__` (line 151) `def __init__(self, modes)`
- `forward` (line 155) `def forward(self, Kfreq)`

#### `inference.py`
**Path:** `topogpt3/inference.py`
**File Doc:** *TopoGPT3 inference engine.  Production-grade autoregressive code completion pipeline for TopoGPT3 checkpoints. Loads weights from safetensors, aligns the underlying TopoGPT2 architecture against the stored tensors, optionally applies the Gauss complex-multiply patch for numerical parity with training, and performs sampling with repetition penalty and top-k filtering.  The pipeline is decomposed into single-responsibility collaborators wired by an orchestrator. All paths, sampling parameters, safety bounds and string identifiers live inside InferenceSettings so that business logic contains no magic numbers or hardcoded constants.*

**Classes:**
- `ScalePreset` (line 32) `class ScalePreset` - *Immutable architecture preset for a named model scale.*
- `InferenceSettings` (line 42) `class InferenceSettings` - *Centralized configuration container for the inference pipeline.

Every value consumed downstream resides here. Adding a new tunable means
extending this class; no other module should embed literals.*
- `InferenceLoggerFactory` (line 160) `class InferenceLoggerFactory` - *Builds a stdout-attached logger from inference settings.*
- `SecurePathResolver` (line 180) `class SecurePathResolver` - *Resolves filesystem paths while rejecting traversal outside their root.*
- `SourceModuleLoader` (line 214) `class SourceModuleLoader` - *Resolves the TopoGPT3 runtime module via the package import system.*
- `CheckpointPaths` (line 230) `class CheckpointPaths` - *Computes and validates checkpoint file paths under a single root.*
- `WeightShapeProbe` (line 276) `class WeightShapeProbe` - *Reads tensor metadata from safetensors to infer architecture details.*
- `TopoGPT2ConfigAligner` (line 320) `class TopoGPT2ConfigAligner` - *Builds a TopoGPT2Config matching the loaded checkpoint and tokenizer.*
- `TokenizerFactory` (line 351) `class TokenizerFactory` - *Builds a BPETokenizer instance using the configured encoding.*
- `GaussPatchApplier` (line 364) `class GaussPatchApplier` - *Applies the idempotent Gauss complex-multiply patch when enabled.*
- `ModelAssembler` (line 382) `class ModelAssembler` - *Instantiates the model and loads weights from safetensors.*
- `SeedSynchronizer` (line 419) `class SeedSynchronizer` - *Applies deterministic seeds across torch, CUDA and the model package.*
- `SamplingPolicy` (line 443) `class SamplingPolicy` - *Immutable sampling parameters consumed by the generation engine.*
- `GenerationReport` (line 463) `class GenerationReport` - *Quantitative summary of a single generation call.*
- `GenerationEngine` (line 477) `class GenerationEngine` - *Runs autoregressive sampling against a loaded model and tokenizer.*
- `ResultRenderer` (line 535) `class ResultRenderer` - *Prints a GenerationReport to stdout using settings-defined formatting.*
- `InferencePipeline` (line 564) `class InferencePipeline` - *Orchestrator wiring loader, builder, engine and renderer.*
- `CliArgumentParser` (line 617) `class CliArgumentParser` - *Translates command-line arguments into an InferenceSettings instance.*

**Methods:**
- `main` (line 723) `def main(argv)` - *CLI entry point. Returns a process exit code.*
- `scale_presets` (line 103) `def scale_presets()` - *Return the architecture preset table indexed by scale name.*
- `preset` (line 118) `def preset(self)` - *Return the resolved preset for the configured model scale.*
- `validate` (line 128) `def validate(self)` - *Raise ValueError if any setting falls outside its safety bounds.*
- `build` (line 164) `def build(settings)` - *Return a configured Logger with a single deduplicated stdout handler.*
- `resolve_under` (line 184) `def resolve_under(root)` - *Join `parts` under `root` and return the canonical resolved path.

Raises ValueError if the resolved path escapes `root`.*
- `require_existing_file` (line 200) `def require_existing_file(path, expected_suffix)` - *Validate `path` points to an existing regular file with the expected suffix.*
- `__init__` (line 217) `def __init__(self, settings, logger)`
- `load` (line 221) `def load(self)` - *Return the topogpt3.train module which re-exports model symbols.*
- `__init__` (line 233) `def __init__(self, settings)`
- `slot_dir` (line 241) `def slot_dir(self)` - *Directory holding the active checkpoint slot.*
- `model_file` (line 245) `def model_file(self)` - *Resolved path to the safetensors weights file inside the slot.*
- `state_file` (line 251) `def state_file(self)` - *Resolved path to the JSON training-state file inside the slot.*
- `assert_ready` (line 257) `def assert_ready(self)` - *Verify weights exist and the on-disk size lies within safety bounds.*
- `__init__` (line 279) `def __init__(self, settings, logger)`
- `detect_n_kv_heads` (line 283) `def detect_n_kv_heads(self, weights_path, d_model, n_heads)` - *Recover N_KV_HEADS used at training by inspecting the k_proj shape.

Returns None when the probe key is absent, signalling the caller to
fall back to scale defaults rather than guess.*
- `__init__` (line 323) `def __init__(self, settings, source_module, logger)`
- `build` (line 329) `def build(self, n_kv_heads, vocab_size)` - *Return a TopoGPT2Config dataclass ready to instantiate the model.*
- `__init__` (line 354) `def __init__(self, settings, source_module)`
- `build` (line 358) `def build(self)` - *Return an instance of BPETokenizer bound to the configured encoding.*
- `__init__` (line 367) `def __init__(self, settings, source_module, logger)`
- `apply_if_enabled` (line 373) `def apply_if_enabled(self)` - *Patch QuaternionSpectralLayer to use the 3-multiply Gauss contract.*
- `__init__` (line 385) `def __init__(self, settings, source_module, logger)`
- `assemble` (line 391) `def assemble(self, aligned_cfg, paths)` - *Build the TopoGPT2 graph, load weights into it, and return it in eval mode.*
- `__init__` (line 422) `def __init__(self, settings, source_module, logger)`
- `apply` (line 428) `def apply(self)` - *Seed all relevant RNGs using the model package helper when available.*
- `from_settings` (line 452) `def from_settings(cls, settings)` - *Construct a SamplingPolicy from inference settings.*
- `tokens_per_second` (line 472) `def tokens_per_second(self, elapsed_floor)` - *Return throughput in tokens/sec, clamped to avoid divide-by-zero.*
- `__init__` (line 480) `def __init__(self, settings, logger)`
- `run` (line 485) `def run(self, model, tokenizer, prompt, policy)` - *Generate a completion for `prompt` and return a GenerationReport.*
- `__init__` (line 538) `def __init__(self, settings, logger)`
- `render` (line 542) `def render(self, report)` - *Emit a banner with prompt and completion, plus a throughput log line.*
- `__init__` (line 567) `def __init__(self, settings, logger)`
- `execute` (line 573) `def execute(self)` - *Run the full inference pipeline end-to-end and return the report.*
- `build_parser` (line 621) `def build_parser()` - *Return the configured argparse.ArgumentParser.*
- `parse` (line 700) `def parse(argv)` - *Parse `argv` (or sys.argv) and return a populated InferenceSettings.*

#### `inference_hrm.py`
**Path:** `topogpt3/inference_hrm.py`
**File Doc:** *TopoGPT3.1: Hierarchical Recursive Reasoning Inference Engine.  This module extends TopoGPT3 with a parameter-free hierarchical recursive reasoning pipeline inspired by:  * Hierarchical Reasoning Model (HRM), Sapient Intelligence: a biologically motivated two-speed architecture with a slow high-level loop and a fast low-level loop. * Tiny Recursive Model (TRM) and Generative Recursive Reasoning Models (GRAM): latent-space recurrence that iterates token vectors until they reach an attractor before projecting them outward.  The pipeline is intentionally built so that the underlying TopoGPT2 weight matrices remain bit-identical to those produced by the TopoGPT3 trainer. No new learnable parameters are introduced. The pretrained transformer layers are repurposed as the recurrent step function of a hierarchical fixed-point iteration whose halting condition is the empirical stabilization of the latent state.  The high-level slow state is persisted across multiple emitted tokens to achieve sparse temporal reasoning: the full network is iterated only at configurable intervals, while a short suffix of layers refines the low-level state at every emitted token.  All configurable values reside in dedicated configuration dataclasses; no magic numbers or hardcoded constants are embedded in business logic. Path resolution rejects traversal escapes. State dict loading defers strictness to settings, so an architecturally aligned TopoGPT3 checkpoint loads unchanged.*

**Classes:**
- `ScalePreset` (line 54) `class ScalePreset` - *Immutable architecture preset for a named model scale.*
- `RecursiveReasoningConfig` (line 64) `class RecursiveReasoningConfig` - *Hyperparameters governing the hierarchical recursive thinking loop.

The semantics follow the HRM and GRAM literature, adapted to operate
safely with zero additional learnable parameters on a model that was
not trained with recurrence in its computational graph. The reasoner
performs damped fixed-point iteration entirely in the residual-stream
space produced by the baseline forward pass; deep activations are never
fed back into the token-embedding-input layers, preserving the trained
activation distribution at every layer boundary.

Attributes:
    enabled: master switch; when False the pipeline degrades to the
        standard non-recursive autoregressive loop.
    max_high_level_iters: maximum slow-loop iterations per emitted token.
        Each iteration applies a deeper trailing window of layers.
    max_low_level_iters: maximum fast-loop iterations per high-level step.
        Each iteration applies the short trailing window of layers.
    low_level_window: number of trailing transformer layers iterated by
        the low-level fast loop.
    high_level_window: number of trailing transformer layers iterated by
        the high-level slow loop. Should be greater than or equal to
        low_level_window so the hierarchy matches the HRM coarse/fine
        split.
    low_level_step: damping coefficient in [0, 1] for the low-level
        update rule z <- z + step * (window(z) - z).
    high_level_step: damping coefficient for the high-level update.
    attractor_low_epsilon: relative L2 change threshold that declares the
        low-level state converged.
    attractor_high_epsilon: relative L2 change threshold that declares the
        high-level state converged.
    high_level_persist_tokens: tokens during which the refinement vector
        is reused as a warm start before being re-initialized to zero.
        This is the sparse temporal-memory dimension.
    cache_warm_start_weight: scalar in [0, 1] applied to the cached
        refinement before warm-starting the next token's iteration.
    max_drift_relative: relative L2 distance ceiling between the iterated
        latent and the baseline latent; exceeding it triggers a reset to
        the baseline state and aborts thinking for the current token.
    latent_change_eps: floor used in the denominator of relative change
        computations to avoid division by zero.
    safety_max_total_iterations: hard cap on total layer invocations per
        emitted token regardless of configured iters.
    minimum_low_level_iters: floor on low-level iterations before
        convergence checks may halt the loop.
    minimum_high_level_iters: floor on high-level iterations before
        convergence checks may halt the loop.
    diagnostic_logging: when True, emits per-token iteration statistics.*
- `HRMInferenceSettings` (line 134) `class HRMInferenceSettings` - *Centralized configuration for the TopoGPT3.1 inference pipeline.

Every value consumed downstream resides here. Extending the pipeline with
a new tunable means extending this dataclass; no other module should
embed literals.*
- `HRMLoggerFactory` (line 343) `class HRMLoggerFactory` - *Builds a stdout-attached logger from inference settings.*
- `SecurePathResolver` (line 363) `class SecurePathResolver` - *Resolves filesystem paths while rejecting traversal outside their root.*
- `SourceModuleLoader` (line 397) `class SourceModuleLoader` - *Resolves the TopoGPT3 runtime module via the package import system.*
- `CheckpointPaths` (line 413) `class CheckpointPaths` - *Computes and validates checkpoint file paths under a single root.*
- `WeightShapeProbe` (line 459) `class WeightShapeProbe` - *Reads tensor metadata from safetensors to infer architecture details.*
- `TopoGPT2ConfigAligner` (line 502) `class TopoGPT2ConfigAligner` - *Builds a TopoGPT2Config matching the loaded checkpoint and tokenizer.*
- `TokenizerFactory` (line 533) `class TokenizerFactory` - *Builds a BPETokenizer instance using the configured encoding.*
- `GaussPatchApplier` (line 546) `class GaussPatchApplier` - *Applies the idempotent Gauss complex-multiply patch when enabled.*
- `ModelAssembler` (line 564) `class ModelAssembler` - *Instantiates the model and loads weights from safetensors.*
- `SeedSynchronizer` (line 601) `class SeedSynchronizer` - *Applies deterministic seeds across torch, CUDA and the model package.*
- `LatentChangeMetric` (line 624) `class LatentChangeMetric` - *Computes the relative L2 distance between two latent tensors.*
- `ReasoningIterationStats` (line 648) `class ReasoningIterationStats` - *Aggregated counters describing a single token's reasoning episode.*
- `GenerationReasoningSummary` (line 661) `class GenerationReasoningSummary` - *Aggregated statistics over the full generation episode.*
- `SparseHighLevelStateCache` (line 684) `class SparseHighLevelStateCache` - *Persists the high-level latent state across consecutive emitted tokens.

The cache is reset whenever its age in tokens reaches the configured
persistence horizon, at which point the next reasoning episode begins
with a zero high-level state. This is the temporal-sparsity mechanism:
expensive full-stack passes are amortized across multiple emissions.*
- `HierarchicalRecursiveReasoner` (line 727) `class HierarchicalRecursiveReasoner` - *Parameter-free hierarchical recursive reasoning over a trained stack.

The reasoner does not own any learnable parameters. It treats the trained
TopoGPT2 transformer layers as a deterministic recurrent step function
and composes them into a two-speed damped fixed-point iteration that
mirrors HRM, while never violating the activation distribution the
trained layers expect.

Algorithm per emitted token:

    1. Run the standard full forward pass once to obtain the baseline
       residual-stream latent z_base and the per-layer kv caches that
       will cross the token boundary. z_base is the trained model's
       native answer for this position.
    2. If recursion is disabled or both iteration budgets are zero,
       return z_base unchanged.
    3. Optionally warm-start z by adding a fraction of the cached
       refinement vector from previous tokens (sparse temporal memory).
    4. Hierarchical refinement, all in residual-stream space:
          for h_step in range(max_high_level_iters):
              for l_step in range(max_low_level_iters):
                  z <- z + low_level_step * (W_low(z) - z)
              z <- z + high_level_step * (W_high(z) - z)
       where W_low and W_high are the last low_level_window and
       high_level_window trained layers respectively, invoked with the
       prefix kv cache treated as immutable. Each update is damped, so
       layer inputs remain close to the trained residual-stream
       distribution.
    5. Hard divergence guard: if the iterated latent drifts farther
       from the baseline than max_drift_relative, reset to the baseline
       and abort thinking for this token. This eliminates the
       catastrophic-attractor failure mode without retraining.
    6. Attractor halting per loop, plus a global cap on total layer
       invocations.

The cached refinement returned to the sparse cache is z_final - z_base,
a small residual-stream displacement that persists across configurable
horizons to amortize thinking effort over multiple tokens.*
- `LogitsSampler` (line 935) `class LogitsSampler` - *Applies temperature, repetition penalty, top-k filtering and multinomial draw.*
- `SamplingPolicy` (line 963) `class SamplingPolicy` - *Immutable sampling parameters consumed by the generation engine.*
- `GenerationReport` (line 985) `class GenerationReport` - *Quantitative summary of a single generation call.*
- `HRMGenerationEngine` (line 1000) `class HRMGenerationEngine` - *Runs autoregressive sampling driven by hierarchical recursive reasoning.

The engine reimplements the prompt encoding and token emission loop so
that the per-token latent state can be intercepted before final norm and
LM-head projection. The intercepted state is handed to a
HierarchicalRecursiveReasoner, which iterates the trained layer stack in
a two-speed loop until the attractor is reached. The final stabilized
latent is then projected to logits and sampled in the standard fashion.*
- `ResultRenderer` (line 1189) `class ResultRenderer` - *Prints a GenerationReport to stdout using settings-defined formatting.*
- `HRMInferencePipeline` (line 1230) `class HRMInferencePipeline` - *Orchestrator wiring loader, builder, reasoner, engine and renderer.*
- `CliArgumentParser` (line 1283) `class CliArgumentParser` - *Translates command-line arguments into an HRMInferenceSettings instance.*

**Methods:**
- `main` (line 1494) `def main(argv)` - *CLI entry point. Returns a process exit code.*
- `scale_presets` (line 221) `def scale_presets()` - *Return the architecture preset table indexed by scale name.*
- `preset` (line 234) `def preset(self)` - *Return the resolved preset for the configured model scale.*
- `validate` (line 244) `def validate(self)` - *Raise ValueError if any setting falls outside its safety bounds.*
- `build` (line 347) `def build(settings)` - *Return a configured Logger with a single deduplicated stdout handler.*
- `resolve_under` (line 367) `def resolve_under(root)` - *Join parts under root and return the canonical resolved path.

Raises ValueError if the resolved path escapes root.*
- `require_existing_file` (line 383) `def require_existing_file(path, expected_suffix)` - *Validate path points to an existing regular file with the expected suffix.*
- `__init__` (line 400) `def __init__(self, settings, logger)`
- `load` (line 404) `def load(self)` - *Return the topogpt3.train module which re-exports model symbols.*
- `__init__` (line 416) `def __init__(self, settings)`
- `slot_dir` (line 424) `def slot_dir(self)` - *Directory holding the active checkpoint slot.*
- `model_file` (line 428) `def model_file(self)` - *Resolved path to the safetensors weights file inside the slot.*
- `state_file` (line 434) `def state_file(self)` - *Resolved path to the JSON training-state file inside the slot.*
- `assert_ready` (line 440) `def assert_ready(self)` - *Verify weights exist and the on-disk size lies within safety bounds.*
- `__init__` (line 462) `def __init__(self, settings, logger)`
- `detect_n_kv_heads` (line 466) `def detect_n_kv_heads(self, weights_path, d_model, n_heads)` - *Recover N_KV_HEADS used at training by inspecting the k_proj shape.

Returns None when the probe key is absent, signalling the caller to
fall back to scale defaults rather than guess.*
- `__init__` (line 505) `def __init__(self, settings, source_module, logger)`
- `build` (line 511) `def build(self, n_kv_heads, vocab_size)` - *Return a TopoGPT2Config dataclass ready to instantiate the model.*
- `__init__` (line 536) `def __init__(self, settings, source_module)`
- `build` (line 540) `def build(self)` - *Return an instance of BPETokenizer bound to the configured encoding.*
- `__init__` (line 549) `def __init__(self, settings, source_module, logger)`
- `apply_if_enabled` (line 555) `def apply_if_enabled(self)` - *Patch QuaternionSpectralLayer to use the 3-multiply Gauss contract.*
- `__init__` (line 567) `def __init__(self, settings, source_module, logger)`
- `assemble` (line 573) `def assemble(self, aligned_cfg, paths)` - *Build the TopoGPT2 graph, load weights into it, and return it in eval mode.*
- `__init__` (line 604) `def __init__(self, settings, source_module, logger)`
- `apply` (line 610) `def apply(self)` - *Seed all relevant RNGs using the model package helper when available.*
- `__init__` (line 627) `def __init__(self, epsilon_floor)`
- `relative_change` (line 632) `def relative_change(self, current, previous)` - *Return ||current - previous|| / max(||previous||, epsilon_floor).*
- `absorb` (line 671) `def absorb(self, sample)` - *Fold a per-token sample into the running totals.*
- `__init__` (line 693) `def __init__(self, persist_tokens)`
- `get_or_init` (line 700) `def get_or_init(self, reference)` - *Return the cached high-level state or a zeroed one when stale.

The boolean flag indicates whether the returned tensor came from a
live cache hit (True) or a fresh zero initialization (False).*
- `commit` (line 716) `def commit(self, new_state)` - *Store a fresh high-level state and increment the cache age.*
- `invalidate` (line 721) `def invalidate(self)` - *Drop any cached state and reset the age counter.*
- `__init__` (line 768) `def __init__(self, layers, final_norm, reasoning_config, logger)`
- `num_layers` (line 789) `def num_layers(self)` - *Return the number of trained transformer layers.*
- `_full_pass` (line 793) `def _full_pass(self, z_in, base_kvs)` - *Forward z_in through every layer using base_kvs as immutable prefix cache.

Returns the layer-stack output and the freshly produced per-layer kv
caches that incorporate the K and V derived from z_in.*
- `_window_pass` (line 808) `def _window_pass(self, z_in, base_kvs, window)` - *Forward z_in through the trailing `window` layers only.

The per-layer kv caches produced during this read-only pass are
discarded; only the baseline pass's committed kvs cross the token
boundary, preserving cache consistency across thinking iterations.*
- `reason` (line 827) `def reason(self, z_initial, base_kvs, cached_refinement)` - *Run hierarchical recursive thinking for a single emission step.

Args:
    z_initial: token embedding of the new position, shape [B, 1, D].
    base_kvs: per-layer kv cache for all previously emitted tokens,
        treated as immutable during thinking iterations.
    cached_refinement: persistent refinement displacement from prior
        tokens, or None to skip the warm start.

Returns:
    A tuple (z_final, committed_kvs, refinement_for_cache, stats):
        z_final is the latent state about to enter the final norm
        and lm head; committed_kvs is the new per-layer kv cache
        including this token's K and V from the baseline pass;
        refinement_for_cache is z_final - z_baseline, to be
        persisted across tokens; stats holds the loop counters.*
- `__init__` (line 938) `def __init__(self, logger)`
- `sample` (line 941) `def sample(self, logits, token_history, temperature, top_k, repetition_penalty)` - *Return a sampled token id tensor of shape [B, 1] from raw logits [B, V].*
- `from_settings` (line 973) `def from_settings(cls, settings)` - *Construct a SamplingPolicy from inference settings.*
- `tokens_per_second` (line 995) `def tokens_per_second(self, elapsed_floor)` - *Return throughput in tokens/sec, clamped to avoid divide-by-zero.*
- `__init__` (line 1011) `def __init__(self, settings, logger)`
- `_encode_prompt` (line 1016) `def _encode_prompt(self, model, prompt_ids)` - *Run the prompt through the full stack once, returning the final
hidden state of the last position, the per-layer base kv caches that
cover all prompt tokens except the last one, and the embedding of the
last prompt token as the seed for the first reasoning episode.*
- `run` (line 1048) `def run(self, model, tokenizer, prompt, policy)` - *Generate a completion for prompt and return a GenerationReport.*
- `__init__` (line 1192) `def __init__(self, settings, logger)`
- `render` (line 1196) `def render(self, report)` - *Emit a banner with prompt, completion, throughput and reasoning stats.*
- `__init__` (line 1233) `def __init__(self, settings, logger)`
- `execute` (line 1239) `def execute(self)` - *Run the full inference pipeline end-to-end and return the report.*
- `build_parser` (line 1287) `def build_parser()` - *Return the configured argparse.ArgumentParser.*
- `parse` (line 1448) `def parse(argv)` - *Parse argv (or sys.argv) and return a populated HRMInferenceSettings.*

#### `jlens.py`
**Path:** `topogpt3/jlens.py`

**Classes:**
- `TopoGPT3JLensFitConfig` (line 37) `class TopoGPT3JLensFitConfig` - *Centralized configuration for Jacobian lens fitting.

Every value consumed downstream resides here. Adding a new tunable means
extending this class; no other module should embed literals.*
- `TopoGPT3JLensAppConfig` (line 56) `class TopoGPT3JLensAppConfig` - *Centralized configuration for Jacobian lens application.

Every value consumed downstream resides here. Adding a new tunable means
extending this class; no other module should embed literals.*
- `ActivationRecorder` (line 69) `class ActivationRecorder` - *Captures residual-stream tensors at the given block indices.

Registers a forward hook on each requested block on ``__enter__`` and
removes them on ``__exit__``. On the next forward pass each block's output
is stored in ``activations``, keyed by block index. Stored tensors are
not detached, so they can be passed straight to ``torch.autograd.grad``.

Args:
    blocks: The sequence of residual blocks (e.g. ``model.layers``).
    at: Block indices to record at.
    start_graph_at: If given, the captured tensor at this index is marked
        ``requires_grad_(True)`` before downstream blocks see it. When the
        model's parameters all have ``requires_grad=False``, this makes the
        captured residual the leaf that roots the autograd graph, so the
        retained graph spans only this block onward.*
- `JacobianLens` (line 459) `class JacobianLens` - *A fitted Jacobian lens: per-layer ``J_l`` matrices and the readout method.

Attributes:
    jacobians: ``{layer_index: Tensor[d_model, d_model]}``. Each ``J_l``
        maps the residual at layer ``l`` into the final-layer basis.
    source_layers: Sorted list of fitted layer indices.
    n_prompts: Number of prompts the lens was averaged over.
    d_model: Residual-stream width.*
- `SliceData` (line 664) `class SliceData` - *Text-format slice data: top-K token predictions per (position, layer).

``layers`` always includes the model's final layer (the actual model
output) so divergences from lens-transported earlier layers are visible.

Attributes:
    seq_len: Number of token positions in the slice.
    layers: Layer indices shown (includes final layer).
    prompt: The input prompt text.
    input_ids: Tensor ``[1, seq_len]`` of token IDs.
    token_strs: Decoded strings for each token position.
    top_ids: ``[seq_len, n_layers, top_n]`` top token IDs per cell.
    top_probs: ``[seq_len, n_layers, top_n]`` softmax probabilities.
    top_token_strs: ``[seq_len, n_layers, top_n]`` decoded token strings
        for each prediction. Empty string if tokenizer was unavailable.*

**Methods:**
- `valid_position_mask` (line 132) `def valid_position_mask(seq_len)` - *Boolean mask over sequence positions to include in the Jacobian average.

Early positions are dominated by attention-sink behaviour and the final
position has no next-token target, so both are excluded.

Args:
    seq_len: Length of the tokenized prompt.
    skip_first: Number of leading positions to exclude.

Returns:
    Boolean tensor of shape ``[seq_len]``.

Raises:
    ValueError: If ``skip_first`` is negative or the prompt is too short to
        leave any valid positions.*
- `_check_layer_indices` (line 162) `def _check_layer_indices(source_layers, target_layer, n_layers)` - *Resolve None/negative layer indices, bounds-check, enforce source < target.*
- `jacobian_for_prompt` (line 187) `def jacobian_for_prompt(model, prompt, source_layers)` - *Compute the per-layer Jacobian estimator ``J_l`` for one prompt.

Runs one forward pass on the prompt replicated ``dim_batch`` times along
the batch axis, retains the graph, then runs ``ceil(d_model / dim_batch)``
backward passes against it. Each backward computes ``dim_batch`` rows of
``J_l`` at once: batch element ``b`` carries a one-hot cotangent at output
dimension ``dim_start + b``, set at every valid target position.

Args:
    model: The model to compute Jacobians for.
    prompt: Input text.
    source_layers: Layer indices ``l`` to compute ``J_l`` at.
    target_layer: Layer to take gradients with respect to. Defaults to the
        final layer; negative indices count from the end.
    dim_batch: Output dimensions computed per backward pass.
    max_seq_len: Truncate the prompt to this many tokens.
    skip_first: Leading positions to exclude.

Returns:
    ``(jacobians, seq_len, n_valid_positions)``. ``jacobians`` maps each
    source layer to a ``[d_model, d_model]`` fp32 CPU tensor.*
- `_atomic_save` (line 283) `def _atomic_save(obj, path)` - *``torch.save`` to a temp file then ``os.replace`` so a crash never
leaves a half-written checkpoint.*
- `fit` (line 291) `def fit(model, prompts)` - *Fit ``J_l`` over a list of prompts and return a JacobianLens.

Per-prompt Jacobians from ``jacobian_for_prompt`` are accumulated as a
running mean. If ``checkpoint_path`` is set, the running sum is written
every ``checkpoint_every`` prompts (atomic) and resumed from on restart.

Args:
    model: The model to fit on.
    prompts: Text prompts to average over.
    source_layers: Layers to fit at. Defaults to every layer below
        ``target_layer``; negative indices count from the end.
    target_layer: See ``jacobian_for_prompt``.
    dim_batch: See ``jacobian_for_prompt``.
    max_seq_len: Truncate each prompt to this many tokens.
    skip_first: See ``jacobian_for_prompt``.
    checkpoint_path: If set, write a resumable checkpoint here.
    checkpoint_every: Write checkpoint every N prompts (default 1).
    resume: If True and checkpoint_path exists, resume from it.

Returns:
    The fitted JacobianLens.

Raises:
    ValueError: If no prompts are long enough to fit on, or if checkpoint
        settings mismatch.*
- `compute_slice` (line 705) `def compute_slice(model, lens, prompt)` - *Compute a position x layer slice of top-K token predictions.

For each layer in the fitted lens, projects the residual at each position
through the Jacobian into the final-layer basis, then unembeds to get
logits and softmax probabilities. Returns the top-N predicted token IDs
and their probabilities per (position, layer) cell.

Args:
    model: The model to read out from.
    lens: A fitted JacobianLens.
    prompt: Input text.
    top_n: Top tokens to keep per (position, layer) cell.
    max_seq_len: Truncate the prompt to this many tokens.

Returns:
    A SliceData instance with arrays indexed ``[seq_len, n_layers, top_n]``.*
- `text_slice` (line 789) `def text_slice(slice_data, tokenizer, n_cols)` - *Render a SliceData as a readable text table showing decoded words.

For each token position, shows what each layer predicts as the next token.
The first column shows the actual input token; subsequent columns show the
top-1 prediction at each layer with its softmax probability. Token strings
are read from ``slice_data.top_token_strs`` (always populated by
``compute_slice``).

Args:
    slice_data: The slice to render.
    tokenizer: Legacy parameter, ignored. Top token strings are already
        stored in ``slice_data.top_token_strs``.
    n_cols: Number of layer columns to show (default 3).

Returns:
    A multi-line string table.*
- `_demo_jlens` (line 842) `def _demo_jlens()` - *Run a full jacobian lens demo loading real weights from checkpoint.*
- `__init__` (line 87) `def __init__(self, blocks, at)`
- `_make_hook` (line 102) `def _make_hook(self, index)`
- `__enter__` (line 113) `def __enter__(self)`
- `__exit__` (line 126) `def __exit__(self)`
- `write_checkpoint` (line 378) `def write_checkpoint()`
- `__init__` (line 470) `def __init__(self, jacobians)`
- `__repr__` (line 482) `def __repr__(self)`
- `save` (line 489) `def save(self, path)` - *Save to ``path``. Jacobians are stored as ``dtype`` (default fp16).*
- `load` (line 504) `def load(cls, path)` - *Load a lens previously written by ``save``.*
- `from_pretrained` (line 519) `def from_pretrained(cls, name_or_path)` - *Load a lens from a local file, a local directory, or a HuggingFace
Hub ``repo_id``.

``filename`` is the path inside the directory or repo; ignored when
``name_or_path`` is itself a file. ``revision`` selects a Hub branch,
tag, or commit.*
- `merge` (line 543) `def merge(cls, lenses)` - *Combine lenses fitted on disjoint prompt subsets into one
(``n_prompts``-weighted mean of the inputs).

Args:
    lenses: Lenses to merge. Must agree on ``source_layers`` and
        ``d_model``.

Raises:
    ValueError: If ``lenses`` is empty or the inputs disagree on shape.*
- `transport` (line 574) `def transport(self, residual, layer)` - *Map a residual at ``layer`` into the final-layer basis: ``J_l @ h``.

Args:
    residual: Tensor of shape ``[..., d_model]``.
    layer: Source layer index (must be in ``source_layers``).*
- `apply` (line 585) `def apply(self, model, prompt)` - *Run ``model`` on ``prompt`` and return lens logits at ``positions``.

Args:
    model: The model to read out from.
    prompt: Input text.
    layers: Layers to read out at. Defaults to all of
        ``source_layers``. Must be a subset of ``source_layers`` when
        ``use_jacobian`` is True.
    positions: Token positions to read out (Python indexing into the
        sequence; negative indices count from the end). None returns
        every position.
    max_seq_len: Truncate the prompt to this many tokens.
    use_jacobian: If False, skip the ``J_l`` transport (vanilla
        logit-lens baseline).

Returns:
    A triple ``(lens_logits, model_logits, input_ids)``. ``lens_logits``
    maps each requested layer to a ``[n_positions, vocab_size]`` tensor;
    ``model_logits`` is the model's actual final-layer logits at the
    same positions (same shape).

Raises:
    ValueError: If any requested layer is out of range for the model,
        or (with use_jacobian) not in source_layers.*
- `__post_init__` (line 692) `def __post_init__(self)`
- `hook` (line 105) `def hook(module, inputs, output)`
- `select` (line 646) `def select(layer)`

#### `lens_model.py`
**Path:** `topogpt3/lens_model.py`

**Classes:**
- `LensModel` (line 23) `class LensModel(Protocol)` - *What the lens needs from a model.

Attributes:
    n_layers: Number of residual blocks.
    d_model: Residual-stream width.
    layers: The residual blocks, indexable by integer; what
        ActivationRecorder hooks.
    tokenizer: Tokenizer used by the visualisation helpers; must provide
        ``decode(token_ids) -> str``. Fitting and apply() never touch it.*
- `TopoGPT3LensConfig` (line 59) `class TopoGPT3LensConfig` - *Centralized configuration for the TopoGPT3 lens model adapter.

Every value consumed downstream resides here. Adding a new tunable means
extending this class; no other module should embed literals.*
- `_TopoGPT3ResidualForward` (line 142) `class _TopoGPT3ResidualForward(Module)` - *Runs the residual block stack only (no final norm, no LM head).

This is the forward subgraph that ActivationRecorder hooks capture.
Extracted from TopoGPT2.forward() to expose the residual stream for
Jacobian lens fitting and application.*
- `TopoGPT3LensModel` (line 161) `class TopoGPT3LensModel(Module)` - *LensModel adapter over a loaded TopoGPT2 model.

Wraps a TopoGPT2 instance and implements the LensModel protocol for use
with ActivationRecorder, JacobianLens fitting, and apply().

The adapter owns no parameters --- all weights live in the wrapped model.
Call ``.eval()`` and set ``requires_grad_(False)`` on the wrapped model
before fitting.*
- `TinyDecoder` (line 306) `class TinyDecoder(Module)` - *A tiny CPU-only decoder for end-to-end tests.

Implements the LensModel protocol indirectly (wrapped by
TopoGPT3LensModel). Residual blocks are ``h + 0.1 * linear(h)``:
the small gain keeps the Jacobian well-conditioned so the late-layer
``diag(J) ~= 1`` property holds.*
- `_ResidualBlock` (line 359) `class _ResidualBlock(Module)`

**Methods:**
- `encode` (line 40) `def encode(self, text)` - *Tokenize ``text`` to ``input_ids`` of shape ``[1, seq_len]`` on the
model's input device.*
- `forward` (line 45) `def forward(self, input_ids)` - *Run the residual stack on ``input_ids`` (no LM head). Must build an
autograd graph through layers when grad is enabled, and must be
deterministic across batch elements (eval mode, dropout off) --- the
fitting estimator replicates the prompt along the batch axis.*
- `unembed` (line 52) `def unembed(self, residual)` - *Map a residual-stream tensor ``[..., d_model]`` to logits
``[..., vocab_size]`` (final norm + LM head).*
- `from_topogpt2_config` (line 84) `def from_topogpt2_config(cls, cfg)` - *Construct a lens config from a TopoGPT2Config dataclass.*
- `probe_checkpoint` (line 104) `def probe_checkpoint(cls, checkpoint_dir)` - *Probe a checkpoint directory and infer lens config from state.json.

Args:
    checkpoint_dir: Path to the checkpoint slot directory.
    state_filename: JSON file containing training config.

Returns:
    A TopoGPT3LensConfig matching the checkpoint.

Raises:
    FileNotFoundError: If state.json is missing.
    ValueError: If required fields are absent from the state.*
- `__init__` (line 150) `def __init__(self, model)`
- `forward` (line 154) `def forward(self, input_ids)`
- `__init__` (line 172) `def __init__(self, model, tokenizer)`
- `n_layers` (line 184) `def n_layers(self)`
- `d_model` (line 188) `def d_model(self)`
- `layers` (line 192) `def layers(self)`
- `tokenizer` (line 196) `def tokenizer(self)`
- `tokenizer` (line 200) `def tokenizer(self, tok)`
- `input_device` (line 204) `def input_device(self)`
- `input_device` (line 210) `def input_device(self, device)`
- `encode` (line 213) `def encode(self, text)` - *Tokenize text to input_ids of shape ``[1, seq_len]``.

Uses BPETokenizer if available, otherwise falls back to a byte-level
encoding compatible with GPT-2 BPE tokenization.*
- `forward` (line 228) `def forward(self, input_ids)` - *Run the residual stack on ``input_ids``.

Returns hidden states of shape ``[batch, seq_len, d_model]``
(pre-final-norm, pre-LM-head). The autograd graph is retained through
all layers when grad is enabled.*
- `unembed` (line 237) `def unembed(self, residual)` - *Map residual ``[..., d_model]`` to logits ``[..., vocab_size]``.

Applies the model's final norm and LM head projection.*
- `from_checkpoint` (line 246) `def from_checkpoint(cls, checkpoint_dir)` - *Build a TopoGPT3LensModel from a checkpoint directory.

Probes state.json for configuration, instantiates the model, loads
safetensors weights, and wraps the result.

Args:
    checkpoint_dir: Path to the checkpoint slot directory.
    device: Target device. Defaults to cuda if available else cpu.
    encoding: Tokenizer encoding name (passed to BPETokenizer).
    strict: Whether to enforce strict state dict loading.

Returns:
    A TopoGPT3LensModel in eval mode with requires_grad_(False).

Raises:
    FileNotFoundError: If model.safetensors or state.json is missing.*
- `__init__` (line 315) `def __init__(self, n_layers, d_model, vocab_size, seed)`
- `forward` (line 344) `def forward(self, token_ids, past_kvs)`
- `__init__` (line 360) `def __init__(self, d_model)`
- `forward` (line 366) `def forward(self, x, past_kv)`

#### `merged_config.py`
**Path:** `topogpt3/merged_config.py`
**File Doc:** *TopoMerged: 6-tier curriculum merging TopoGPT3 code tasks with ExploitGym exploit tasks.  Tier structure: Phase 1 — Code foundation (from TopoGPT3): Tier 1: codealpaca            (128 seq, 2 epochs) — short instructions Tier 2: code_feedback         (192 seq, 2 epochs) — step-by-step explanations Tier 3: magicoder_evol        (256 seq, 3 epochs) — complex problems  Phase 2 — Exploit specialization (from ExploitGym): Tier 4: vulnerability_analysis (256 seq, 4 epochs) — analyze CVEs Tier 5: patch_analysis         (512 seq, 3 epochs) — understand patches Tier 6: exploit_development    (768 seq, 4 epochs) — write exploits  The model learns to code first, then applies coding skill to security.  Author: Gris Iscomeback License: GPL v3*

**Classes:**
- `TopoMergedConfig` (line 32) `class TopoMergedConfig` - *Configuration for TopoMerged: 6-tier code+exploit curriculum.*

**Methods:**
- `build_topogpt2_config` (line 139) `def build_topogpt2_config(self, max_seq_len, attn_window)`
- `build_exploit_loader_config` (line 161) `def build_exploit_loader_config(self)`
- `is_code_tier` (line 175) `def is_code_tier(self, tier_index)` - *Returns True if the tier uses TopoGPT3 (HuggingFace) data.*
- `is_exploit_tier` (line 179) `def is_exploit_tier(self, tier_index)` - *Returns True if the tier uses ExploitGym data.*
- `exploit_tier_index` (line 183) `def exploit_tier_index(self, tier_index)` - *Convert global tier index to exploitgym-local index (0-2).*

#### `model.py`
**Path:** `topogpt3/model.py`
**File Doc:** *TopoGPT2: Quaternion-Enhanced Topological Transformer Language Model  Author: Gris Iscomeback Email: grisiscomeback@gmail.com License: GPL v3  Mejoras sobre topogpt.py: - Álgebra de cuaterniones completa (QuaternionLinear, QuaternionSpectralLayer) con producto de Hamilton en el dominio de frecuencia para capturar la espectrografía de los datos con kernels reales e imaginarios cruzados. - SpectralAutoencoder: encoder/decoder espectral que comprime y reconstruye las representaciones en el dominio de frecuencia. - QuaternionTorusBrain VECTORIZADA (sin bucles sobre seq_len): proyección geométrica sobre el toro con asignación blanda usando distancias circulares, message-passing con rotaciones de cuaterniones. - 8 nodos (RADIAL=2 × ANGULAR=4), 4 ángulos, 2 radiales (spec del usuario). - Rotary Position Embeddings (RoPE). - Flash-attention (scaled_dot_product_attention de PyTorch 2.0+). - RMSNorm en lugar de LayerNorm (estilo LLaMA). - Tokenizador BPE via tiktoken (vocab GPT-2, 50k tokens). - Descargador de corpus: TinyStories, WikiText-103, raw file. - Entrenamiento con AMP (mixed precision) + acumulación de gradientes. - Presets de escala: micro, small, medium, gpt2.*

**Classes:**
- `TopoGPT2Config` (line 56) `class TopoGPT2Config` - *Configuración completa para TopoGPT2.*
- `QuaternionOps` (line 216) `class QuaternionOps` - *Operaciones de cuaterniones puras en PyTorch.
Representación: [..., 4]  donde last dim = [w, x, y, z]
q = w + x*i + y*j + z*k*
- `QuaternionLinear` (line 255) `class QuaternionLinear(Module)` - *Capa lineal con pesos cuaterniones.

Implementa la multiplicación W * x en el álgebra de cuaterniones:
- W = Ww + Wx*i + Wy*j + Wz*k  (cuaternión de pesos)
- x = xw + xx*i + xy*j + xz*k  (cuaternión de entrada)
- out = W * x  (producto de Hamilton extendido a vectores)

Parámetros: 4 matrices reales de forma [out_q, in_q]*
- `QuaternionSpectralLayer` (line 300) `class QuaternionSpectralLayer(Module)` - *Convolución espectral 2D con cuaterniones y producto de Hamilton completo.

Operación en dominio de frecuencia:
    P(k) = W(k) ⊗ X(k)  (producto de Hamilton de cuaterniones complejos)

Donde:
    X(k) = FFT2(x) con 4 canales cuaterniones [Xw, Xx, Xy, Xz]
    W(k) = kernel complejo aprendible con componentes [Ww, Wx, Wy, Wz]

Reglas del producto de Hamilton en dominio de frecuencia:
    Pw = Ww·Xw - Wx·Xx - Wy·Xy - Wz·Xz
    Px = Ww·Xx + Wx·Xw + Wy·Xz - Wz·Xy
    Py = Ww·Xy - Wx·Xz + Wy·Xw + Wz·Xx
    Pz = Ww·Xz + Wx·Xy - Wy·Xx + Wz·Xw

Cada Wc es un kernel complejo (partes real e imaginaria independientes).*
- `SpectralAutoencoder` (line 387) `class SpectralAutoencoder(Module)` - *Autoencoder espectral con cuaterniones.

Opera en dos niveles:
1. Espectral 1D sobre el vector de features (FFT sobre dim D_MODEL):
   captura la espectrografía global del embedding.
2. Espectral 2D sobre el grid del toro (QuaternionSpectralLayer):
   captura correlaciones espaciales en la topología.

Devuelve (latent, recon_loss) para regularización.*
- `QuaternionTorusBrain` (line 470) `class QuaternionTorusBrain(Module)` - *Reemplaza el MLP en cada capa del transformer.

Pipeline (completamente vectorizado sobre batch Y secuencia):

1. Flatten: [B, S, D] → [B·S, D]
2. SpectralAutoencoder: filtrado espectral 1D + compresión cuaternión
3. Proyección al toro:
   - Calcula 2 ángulos (phi1, phi2) ∈ [-π, π]²
   - Asignación blanda a los 8 nodos via distancia circular en el toro
4. Construye grid de nodos: [B·S, N_NODES=8, D_MODEL]
5. QuaternionSpectralLayer 2D sobre el grid [B·S, 4*D_QUAT, RADIAL, ANGULAR]
6. Message-passing con rotaciones cuaterniones sobre el grafo toro
7. Readout: atención sobre los 8 nodos → [B·S, D_MODEL]
8. Reshape: [B·S, D] → [B, S, D]*
- `RotaryEmbedding` (line 687) `class RotaryEmbedding(Module)` - *Rotary Position Embeddings (RoPE) - Su et al., 2021.
Codifica la posicion como rotaciones del espacio de atencion,
naturalmente relativas y sin parametros extra.

Las caches _cos/_sin se registran como buffers no-persistentes con
nombres que no colisionan con checkpoints antiguos (que usaban
'cos_cache'/'sin_cache'). Esto permite cambiar MAX_SEQ_LEN sin
errores de shape al cargar checkpoints previos.*
- `RMSNorm` (line 741) `class RMSNorm(Module)` - *Root Mean Square Layer Normalization (sin bias). Más estable que LayerNorm.*
- `SwiGLU` (line 758) `class SwiGLU(Module)` - *SwiGLU: SiLU(gate(x)) * up(x) -> down
Usado en LLaMA 2/3, Qwen, Mistral en lugar de GELU-FFN.
Dimension interna: 8/3 * d_model (convención LLaMA, redondeada a múltiplo de 4).*
- `TopoMoEBrain` (line 787) `class TopoMoEBrain(Module)` - *Mixture of Experts sobre la capa topologica.

Arquitectura (inspirada en DeepSeek-MoE / Mixtral):
  - 1 experto compartido: QuaternionTorusBrain (siempre activo)
  - N_EXPERTS expertos SwiGLU ligeros (activacion esparsa: Top-K por token)
  - Router: Linear(D, N_EXPERTS) + softmax → top-K

Load-balancing loss (auxiliar): penaliza si un experto acapara todos los tokens.
Activa MOE_TOP_K de N_EXPERTS expertos por token.

Sin MoE (MOE_ENABLED=False): se comporta como QuaternionTorusBrain puro.*
- `MultiHeadAttention` (line 892) `class MultiHeadAttention(Module)` - *Multi-head attention con:
- Flash Attention (scaled_dot_product_attention de PyTorch 2.0+)
- Rotary Position Embeddings (RoPE)
- GQA (Grouped Query Attention): N_KV_HEADS < N_HEADS, reduce VRAM de K/V
- KV Cache para inferencia autoregresiva eficiente
- Temperatura termodinámica aprendible*
- `TopoGPT2Layer` (line 996) `class TopoGPT2Layer(Module)` - *Capa del transformer con TopoMoEBrain (TopoBrain + MoE SwiGLU experts).

Esquema pre-norm (estilo LLaMA):
    x = x + Attention_GQA(RMSNorm(x))
    x = x + TopoMoEBrain(RMSNorm(x))*
- `TopoGPT2` (line 1043) `class TopoGPT2(Module)` - *TopoGPT2: Transformer de lenguaje con TopoBrain cuaternión-espectral.

Arquitectura:
    Embedding de tokens + RoPE (en Attention)
    N_LAYERS × TopoGPT2Layer (Attention + QuaternionTorusBrain)
    RMSNorm final
    Proyección a vocabulario (weight-tied con embeddings)*
- `BPETokenizer` (line 1260) `class BPETokenizer` - *Wrapper alrededor de tiktoken (GPT-2 compatible).*
- `FileManifest` (line 1386) `class FileManifest` - *Disk-cached manifest of text files found in a directory tree.*
- `MemmapTokenizer` (line 1453) `class MemmapTokenizer` - *Tokenizes file paths into a memory-mapped numpy array on disk.

Uses incremental file reading and batched writing to avoid loading
all tokens into RAM. Tokens are stored as raw int64 on disk and
accessed via numpy memmap (OS-level paging, near-zero RAM footprint).*
- `MappedTokenDataset` (line 1542) `class MappedTokenDataset(Dataset)` - *Memory-mapped token dataset for sequence-to-sequence LM training.

The token array is backed by a numpy memmap file on disk.
Only accessed slices are paged into RAM by the OS. The .copy()
in __getitem__ ensures the returned torch.Tensor owns its memory,
which is required for DataLoader collation with worker processes.*
- `TextFilter` (line 1576) `class TextFilter` - *Filters low-quality files from the corpus based on multiple heuristics.*
- `CurriculumDataset` (line 1680) `class CurriculumDataset(Dataset)` - *Tiered dataset that exposes short/medium/all files based on line count.

Works as a wrapper around MappedTokenDataset. Provides __getitem__ that
only samples from the active tier, avoiding dataset duplication.*
- `ProgressiveSeqLenTrainer` (line 1744) `class ProgressiveSeqLenTrainer` - *Trainer that dynamically adjusts MAX_SEQ_LEN across training phases.

Phase schedule (configurable):
    phase 0: seq_len=128, epochs=3
    phase 1: seq_len=256, epochs=3
    phase 2: seq_len=512, epochs=4

Each phase rebuilds the DataLoader with the new sequence length.*
- `SpeculativeDecoder` (line 1820) `class SpeculativeDecoder` - *Speculative decoding with a small draft model.

Draft model uses SPEC_DECODE_DRAFT_SCALE (e.g. 'micro').
The draft generates K tokens, then the target model verifies them
in a single forward pass. Accepted tokens are kept; rejected ones
trigger a fallback to the target model sampling.*
- `QuantizedEmbedding` (line 1935) `class QuantizedEmbedding(Module)` - *Wrapper around nn.Embedding that applies dynamic quantization.

Applies int8 quantization to the embedding weight matrix after loading.
Supports both embed (int8) and FFN (int4) quantization modes.*
- `CurriculumTrainer` (line 2003) `class CurriculumTrainer` - *Extends TopoGPT2Trainer with curriculum + progressive seq len support.

Provides:
- Tokens cache for progressive sequence length rebuilding
- Curriculum dataset wrapping (short / medium / all tiers)*
- `CheckpointManager` (line 2155) `class CheckpointManager` - *Gestiona checkpoints de forma acumulativa y segura.

Estructura en disco:
    checkpoints_topogpt2/
      latest/
        model.safetensors   <- pesos del modelo (formato seguro, sin pickle)
        optimizer.pt        <- estado del optimizador (requiere .pt)
        state.json          <- metadatos: epoch, step, historial, config
      best/
        model.safetensors
        state.json
      step_NNNNN/           <- snapshots periodicos (rotados)
        model.safetensors
        optimizer.pt
        state.json

El historial se ACUMULA entre sesiones de entrenamiento: cada --resume
agrega nuevas entradas a train_loss[], val_loss[], etc.*
- `TopoGPT2Trainer` (line 2388) `class TopoGPT2Trainer` - *Entrenador acumulativo y resumible.

Caracteristicas:
- Checkpoint automatico en safetensors cada N minutos + cada epoch
- Historial acumulativo entre sesiones (--resume agrega al historial existente)
- Guarda el mejor modelo en checkpoints/best/ automaticamente
- LR schedule: cosine con warmup relativo a los steps de ESTA sesion
- Mixed Precision (AMP) + acumulacion de gradientes*
- `MechanisticMetrics` (line 2681) `class MechanisticMetrics` - *Calcula todas las metricas del diagrama de fases de Book.md.

Todas las metricas se derivan de cantidades medibles (pesos, gradientes):

delta  (δ): margen de discretizacion.  max|w - round(w)|
            δ≈0 -> cristal;  δ≈0.49 -> vidrio frio
kappa  (κ): numero de condicion de la covarianza del gradiente.
            κ≈1 -> cristalino;  κ>>1 -> amorfo
T_eff:      temperatura efectiva = (lr/2) * Var(gradiente).
            T_eff→0 -> congelado; T_eff alto -> ruidoso
alpha  (α): indice de pureza = -log(δ + ε).
            α=20 -> perfecto; α<1 -> vidrio
berry:      fase de Berry de los kernels espectrales imaginarios.
            |berry|>π/2 con winding≠0 -> insulador topologico
lc:         complejidad local = 1 - similitud coseno promedio entre filas.
sp:         superposicion = correlacion promedio inter-fila de pesos.*
- `Phase0_KernelOptimizer` (line 2918) `class Phase0_KernelOptimizer` - *Encuentra el ratio imaginario/real optimo para los kernels espectrales.

Analogia con main.py: evalua la transicion GOE→GUE en el espacio
de kernels. Un ratio optimo promueve estructura topologica (insulador)
vs estructura amorfa (vidrio).

Metodo: calibra con un mini-batch y mide la varianza del gradiente
en funcion del ratio. Ratios que minimizan la varianza de gradiente
(maxima coherencia espectral) son preferibles.

No entrena: solo inicializa los kernels con distintos ratios y mide.
Tiempo tipico: < 30 segundos.*
- `Phase1_BatchProspector` (line 2993) `class Phase1_BatchProspector` - *Encuentra el batch size optimo testando candidatos con pocos pasos.

De main.py: el batch size regula la temperatura del horno de cristalizacion.
Batch sizes demasiado chicos -> ruido excesivo (vidrio frio).
Batch sizes demasiado grandes -> sin presion annealing (amorfos).
La ventana optima empirica de main.py: [24, 128] para Strassen.

Para LM, testeamos candidatos midiendo:
- delta (δ): velocidad de descenso en prospect_steps pasos
- T_eff: temperatura efectiva del gradiente

Tiempo tipico: < 2 minutos para 3 candidatos × 30 pasos.*
- `Phase2_SeedMiner` (line 3076) `class Phase2_SeedMiner` - *Encuentra semillas prometedoras midiendo la trayectoria de delta.

De main.py: una semilla "buena" muestra delta descendente en los
primeros N pasos (enfriamiento). Una semilla "mala" se estanca en
el plateau vidrioso (~0.49).

Criterio de seleccion:
1. Semillas con delta_velocity < 0 (enfriando) AND kappa bajo.
2. Si no hay, semillas solo enfriando.
3. Fallback: semilla con menor delta final.

Tiempo tipico: < 3 minutos para 5 semillas × 50 pasos.*
- `Phase4_AnnealingRefiner` (line 3158) `class Phase4_AnnealingRefiner` - *Refinamiento post-entrenamiento mediante recocido simulado.

De main.py: despues de que el modelo converge, una fase de annealing
con criterio de aceptacion de Metropolis puede empujar los pesos
hacia estados de menor energia libre (menor delta o mejor val_loss).

Aceptacion de Metropolis:
    si Δloss < 0: siempre acepta (mejora)
    si Δloss >= 0: acepta con prob exp(-Δloss / T)

La temperatura T decae exponencialmente: T(t) = T0 * cooling_rate^t

Al rechazar: restaura el mejor estado conocido.
Si se estanca: perturbacion termica (ruido gaussiano en pesos).

Tiempo: proporcional a refine_epochs (user-controlled).*
- `TopoPhasePipelineV2` (line 3319) `class TopoPhasePipelineV2` - *Pipeline with curriculum learning and progressive sequence length.

Replaces TopoPhasePipeline when --curriculum or --progressive-seq-len is set.
Handles:
- Text quality filtering before tokenization (via TextFilter)
- Curriculum tiers (short/medium/all files)
- Progressive MAX_SEQ_LEN across phases: 128->256->512
- Tokens cached in memory for fast DataLoader rebuilding per phase*
- `TopoPhasePipeline` (line 3448) `class TopoPhasePipeline` - *Orquesta las 5 fases de entrenamiento segun main.py + Book.md.

Fases:
  0  Kernel ratio optimization  (GOE-GUE spectral calibration)
  1  Batch size prospecting      (temperatura del horno de cristalizacion)
  2  Seed mining                 (seleccion de semilla enfriante)
  3  Full training               (entrenamiento principal con metricas)
  4  Annealing refinement        (recocido simulado post-entrenamiento)

Las fases 0-2 son rapidas (prospecting). La fase 3 es el grueso.
La fase 4 es opcional (--refine).

Para no ser prohibitivo:
  --prospect         activa fases 0, 1, 2 antes del entrenamiento
  --refine-epochs N  activa fase 4 con N epocas de annealing
  Sin flags: solo fase 3 (comportamiento original, identico a antes)*

**Methods:**
- `setup_logger` (line 194) `def setup_logger(name, level)`
- `set_seed` (line 204) `def set_seed(seed, device)`
- `build_file_tiers` (line 1718) `def build_file_tiers(paths, short, med)` - *Classify file paths into complexity tiers by line count.

Returns dict: tier -> list of file indices in that tier.
Tier 0 = short (<=short lines), tier 1 = medium, tier 2 = all.*
- `apply_quantization` (line 1975) `def apply_quantization(model, config)` - *Quantize embedding and lm_head layers for reduced VRAM usage.*
- `_tokenize_text_to_memmap` (line 2143) `def _tokenize_text_to_memmap(text, tokenizer, path, max_tokens)` - *Tokenize a single text string and write tokens to disk as raw int64.*
- `main` (line 3570) `def main()`
- `__post_init__` (line 161) `def __post_init__(self)`
- `hamilton_product` (line 224) `def hamilton_product(q1, q2)` - *Producto de Hamilton q1 ⊗ q2. Ambos [..., 4].*
- `normalize` (line 236) `def normalize(q, eps)`
- `conjugate` (line 240) `def conjugate(q)`
- `rotate_vector` (line 245) `def rotate_vector(v, q)` - *Rota vector 3D v por cuaternión unitario q. v:[...,3] q:[...,4]*
- `__init__` (line 267) `def __init__(self, in_features, out_features, bias)`
- `forward` (line 283) `def forward(self, x)` - *x: [..., in_features] → [..., out_features]*
- `__init__` (line 320) `def __init__(self, in_q, out_q, grid_h, grid_w, init_scale)`
- `_kernel` (line 339) `def _kernel(self, c)`
- `_contract` (line 342) `def _contract(self, W, X)` - *Suma sobre canales in_q: Y[b,o,h,w] = Σ_i W[i,o,h,w]·X[b,i,h,w]*
- `forward` (line 346) `def forward(self, x)` - *x: [B, 4*in_q, H, W]  (4 canales cuaterniones sobre grid espacial)
→ [B, 4*out_q, H, W]*
- `__init__` (line 400) `def __init__(self, config)`
- `_filter1d` (line 432) `def _filter1d(self, x, kr, ki)` - *Filtro espectral 1D: x[..., D] → filtrado[..., D]*
- `encode` (line 438) `def encode(self, x)` - *x: [..., D_MODEL] → latent: [..., D_LAT]*
- `decode` (line 443) `def decode(self, z)` - *z: [..., D_LAT] → recon: [..., D_MODEL]*
- `forward` (line 448) `def forward(self, x)` - *Devuelve (latent, recon_loss)*
- `process_torus_grid` (line 455) `def process_torus_grid(self, grid)` - *Procesa el grid del toro con QuaternionSpectralLayer.
grid: [B, 4*D_QUAT, RADIAL, ANGULAR]  →  [B, 4*D_QUAT, RADIAL, ANGULAR]*
- `__init__` (line 488) `def __init__(self, d_model, config)`
- `_build_torus_graph` (line 528) `def _build_torus_graph(self)` - *Construye las aristas del grafo toro 2×4.

Nodos indexados como: node = r * N_ANGULAR + a
  r ∈ [0, RADIAL-1], a ∈ [0, ANGULAR-1]

Aristas angulares: nodo ↔ nodo a la izquierda/derecha (periódico)
Aristas radiales:  nodo ↔ nodo del anillo interior/exterior*
- `_torus_soft_assign` (line 562) `def _torus_soft_assign(self, phi1, phi2)` - *Asignación blanda de tokens a los 8 nodos del toro via distancia circular.

phi1: [BS] ángulo angular ∈ [-π, π]
phi2: [BS] ángulo radial ∈ [-π, π]
→ weights: [BS, N_NODES]  (suma a 1, softmax de distancias negativas)*
- `_message_passing` (line 589) `def _message_passing(self, node_feat)` - *Message-passing VECTORIZADO con rotaciones cuaterniones.
Sin bucles Python: todas las aristas se procesan en paralelo.

node_feat: [BS, N_NODES, D_MODEL]
→ [BS, N_NODES, D_MODEL]*
- `forward` (line 626) `def forward(self, x)` - *x: [B, S, D_MODEL]
→ output: [B, S, D_MODEL], recon_loss: scalar*
- `__init__` (line 699) `def __init__(self, d_head, max_seq_len, base)`
- `_build_cache` (line 705) `def _build_cache(self, seq_len)`
- `_rotate_half` (line 713) `def _rotate_half(self, x)`
- `forward` (line 717) `def forward(self, q, k, seq_len, offset)` - *q, k: [B, n_heads, S_q/S_k, d_head]
offset: posicion inicial (para KV cache: longitud del cache existente)
Aplica posiciones [offset .. offset+S-1] a q y k.*
- `__init__` (line 744) `def __init__(self, d_model, eps)`
- `forward` (line 749) `def forward(self, x)`
- `__init__` (line 765) `def __init__(self, d_model, expansion, dropout)`
- `forward` (line 779) `def forward(self, x)`
- `__init__` (line 802) `def __init__(self, d_model, config)`
- `_route` (line 823) `def _route(self, x)` - *x: [N, D] donde N = B*S (tokens aplanados)
Retorna:
expert_out: [N, D]  suma ponderada de top-K expertos
aux_loss:   escalar  load-balancing loss
Routing vectorizado sin boolean indexing ni sincronizacion CUDA.
Usa dispatch por indices agrupados (estilo Mixtral/DeepSeek) para
compatibilidad total con torch.utils.checkpoint.*
- `forward` (line 865) `def forward(self, x)` - *x: [B, S, D]
→ output: [B, S, D], aux_loss: escalar*
- `__init__` (line 902) `def __init__(self, d_model, n_heads, config)`
- `forward` (line 921) `def forward(self, x, is_causal, past_kv)` - *Args:
    x:        [B, S, D]
    is_causal: usar mascara causal
    past_kv:  (K_cache, V_cache) de pasos anteriores o None
Returns:
    out:      [B, S, D]
    kv_cache: (K, V) completos para cachear en generate()*
- `__init__` (line 1005) `def __init__(self, d_model, n_heads, config)`
- `_forward_impl` (line 1014) `def _forward_impl(self, x, past_kv)`
- `forward` (line 1023) `def forward(self, x, past_kv)` - *Retorna (x_out, aux_loss, kv_cache).
Con gradient checkpointing en training (solo cuando no hay KV cache).*
- `__init__` (line 1054) `def __init__(self, config)`
- `_init_weights` (line 1079) `def _init_weights(self)`
- `forward` (line 1086) `def forward(self, token_ids, past_kvs)` - *token_ids: [B, S]  (enteros)
past_kvs:  lista de (K, V) por capa, o None para entrenamiento
→ logits: [B, S, VOCAB_SIZE], aux_loss: scalar, new_kvs: list[(K,V)]*
- `forward_with_memory` (line 1109) `def forward_with_memory(self, token_ids)` - *Process long sequences with latent memory-token context compression.

Splits `token_ids` [B, S] into segments of size MEMORY_SEGMENT_LEN.
Each segment is processed with N_MEMORY_TOKENS prepended. The output
at memory-token positions after segment k becomes the memory-state
input for segment k+1, compressing all prior context into a fixed-size
latent vector.

Returns (logits [B, S, VOCAB_SIZE], aux_loss).*
- `count_params` (line 1159) `def count_params(self)`
- `generate` (line 1165) `def generate(self, token_ids, max_new_tokens, temperature, top_k, repetition_penalty)` - *Autoregressive generation with KV cache and top-k sampling.

Args:
    token_ids: [B, S_prompt] prompt tokens.
    max_new_tokens: Maximum tokens to generate.
    temperature: Sampling temperature (lower = more deterministic).
    top_k: Top-k filtering (0 = disabled).
    repetition_penalty: Penalty for repeating tokens (>1 = penalize).

Returns:
    [B, S_prompt + generated] full token sequence.*
- `generate_with_continuation` (line 1216) `def generate_with_continuation(self, token_ids, tokenizer, max_new_tokens, temperature, top_k, repetition_penalty, max_continuations, tail_lines)`
- `__init__` (line 1263) `def __init__(self, encoding)`
- `encode` (line 1271) `def encode(self, text)`
- `decode` (line 1274) `def decode(self, tokens)`
- `eot_token` (line 1277) `def eot_token(self)`
- `__init__` (line 1389) `def __init__(self, root, cache_dir, logger)`
- `scan` (line 1396) `def scan(self, force)` - *Walk directory tree collecting text file paths. Cached to disk.*
- `__init__` (line 1463) `def __init__(self, cache_dir, logger)`
- `tokenize` (line 1468) `def tokenize(self, file_paths, tokenizer, cache_key, max_tokens, min_chars)` - *Tokenize all files and return a memory-mapped numpy array.

Args:
    file_paths: List of absolute file paths to tokenize.
    tokenizer: BPE tokenizer instance.
    cache_key: Unique key for caching tokens to disk.
    max_tokens: Maximum number of tokens to produce.
    min_chars: Skip files with fewer characters.

Returns:
    np.ndarray backed by a memmap on disk. Only accessed pages
    are loaded into RAM by the OS virtual memory system.*
- `__init__` (line 1551) `def __init__(self, tokens, seq_len)`
- `__len__` (line 1556) `def __len__(self)`
- `__getitem__` (line 1559) `def __getitem__(self, idx)`
- `__init__` (line 1579) `def __init__(self, config, logger)`
- `_compute_entropy` (line 1588) `def _compute_entropy(self, text)` - *Shannon entropy of byte frequencies (bits per byte).*
- `_has_long_lines` (line 1602) `def _has_long_lines(self, text, threshold)` - *Return True if any line exceeds threshold characters.*
- `_special_token_ratio` (line 1609) `def _special_token_ratio(self, text, tokenizer)` - *Fraction of tokens that are pure whitespace or indentation-only.*
- `_content_hash` (line 1623) `def _content_hash(self, text)`
- `filter_file` (line 1626) `def filter_file(self, path, tokenizer)` - *Read and evaluate a file. Returns text if passed, None if filtered.*
- `report` (line 1666) `def report(self)`
- `__init__` (line 1687) `def __init__(self, tokens, seq_len, file_tiers, active_tier, logger)`
- `_update_len` (line 1697) `def _update_len(self)`
- `set_tier` (line 1703) `def set_tier(self, tier)`
- `__len__` (line 1707) `def __len__(self)`
- `__getitem__` (line 1710) `def __getitem__(self, idx)`
- `__init__` (line 1755) `def __init__(self, base_trainer)`
- `_build_dataloader` (line 1761) `def _build_dataloader(self, dataset, seq_len, batch_size, is_train)`
- `run` (line 1771) `def run(self, train_paths, val_paths, tokenizer, file_tiers, phases)` - *Run training with progressive sequence length across phases.*
- `__init__` (line 1829) `def __init__(self, target_model, config, logger)`
- `_build_draft` (line 1837) `def _build_draft(self)`
- `generate` (line 1851) `def generate(self, token_ids, max_new_tokens, temperature, top_k, repetition_penalty)` - *Autoregressive generation via speculative decoding.

Each round: draft generates K tokens, target verifies all K in
one O(1) forward pass (longest context), then samples the first
rejection from the target.*
- `__init__` (line 1942) `def __init__(self, embed, mode)`
- `forward` (line 1971) `def forward(self, indices)`
- `__init__` (line 2011) `def __init__(self, model, config, tokenizer)`
- `cache_tokens` (line 2018) `def cache_tokens(self, key, tokens)`
- `model` (line 2022) `def model(self)`
- `optimizer` (line 2026) `def optimizer(self)`
- `scaler` (line 2030) `def scaler(self)`
- `amp_dtype` (line 2034) `def amp_dtype(self)`
- `completed_epochs` (line 2038) `def completed_epochs(self)`
- `completed_epochs` (line 2042) `def completed_epochs(self, v)`
- `global_step` (line 2046) `def global_step(self)`
- `global_step` (line 2050) `def global_step(self, v)`
- `best_val_loss` (line 2054) `def best_val_loss(self)`
- `best_val_loss` (line 2058) `def best_val_loss(self, v)`
- `history` (line 2062) `def history(self)`
- `ckpt_mgr` (line 2066) `def ckpt_mgr(self)`
- `resume` (line 2069) `def resume(self)`
- `_current_state` (line 2072) `def _current_state(self)`
- `_cosine_lr` (line 2075) `def _cosine_lr(self)`
- `_set_lr` (line 2078) `def _set_lr(self)`
- `evaluate` (line 2081) `def evaluate(self, dataloader)`
- `_sample_text` (line 2084) `def _sample_text(self)`
- `_progressive_train` (line 2087) `def _progressive_train(self, train_paths, val_paths, tokenizer, phases, memtok)` - *Training loop with progressive sequence length across phases.*
- `train` (line 2126) `def train(self, train_dl, val_dl)`
- `run_curriculum` (line 2129) `def run_curriculum(self, train_paths, val_paths, tokenizer, phases)` - *Top-level entry point: curriculum + progressive seq len.*
- `__init__` (line 2180) `def __init__(self, config, logger)`
- `patch_config_for_resume` (line 2190) `def patch_config_for_resume(self, cfg)` - *Lee el checkpoint 'latest' y ajusta cfg.N_KV_HEADS / cfg.GQA_GROUPS
para que coincidan con la arquitectura guardada.
Necesario cuando el codigo cambio GQA despues de guardar el checkpoint.*
- `_save_model` (line 2219) `def _save_model(self, model, directory)`
- `_load_model` (line 2232) `def _load_model(self, model, directory)`
- `_save_optimizer` (line 2263) `def _save_optimizer(self, optimizer, directory)`
- `_load_optimizer` (line 2266) `def _load_optimizer(self, optimizer, directory, device)`
- `_save_state` (line 2275) `def _save_state(self, state, directory)`
- `_load_state` (line 2280) `def _load_state(self, directory)`
- `should_save` (line 2291) `def should_save(self)`
- `save` (line 2294) `def save(self, model, optimizer, state, is_best)` - *Guarda checkpoint completo.

state debe contener al menos: completed_epochs, global_step,
best_val_loss, history, config.*
- `load_latest` (line 2339) `def load_latest(self, model, optimizer)` - *Carga el ultimo checkpoint guardado.
Devuelve el state dict (vacio si no hay checkpoint).*
- `load_best` (line 2366) `def load_best(self, model)` - *Carga el mejor modelo guardado (solo pesos, sin optimizador).*
- `has_checkpoint` (line 2378) `def has_checkpoint(self)`
- `__init__` (line 2400) `def __init__(self, model, config, tokenizer)`
- `resume` (line 2435) `def resume(self)` - *Carga el ultimo checkpoint disponible.
Restaura: pesos del modelo, estado del optimizador, historial acumulado,
epoch/step completados y mejor val_loss.
Devuelve True si se cargo un checkpoint, False si empieza de cero.*
- `_current_state` (line 2460) `def _current_state(self)` - *Construye el dict de estado para persistir en state.json.*
- `_cosine_lr` (line 2471) `def _cosine_lr(self, step_in_session, total_steps_session)` - *Cosine decay con warmup. El schedule es relativo a la sesion actual.*
- `_set_lr` (line 2479) `def _set_lr(self, lr)`
- `train` (line 2483) `def train(self, train_dl, val_dl)` - *Entrena cfg.EPOCHS epocas adicionales a partir de completed_epochs.
El historial se acumula sobre sesiones previas.*
- `_sample_text` (line 2619) `def _sample_text(self, tokenizer, prompts, max_new, temperature, top_k)` - *Genera una muestra de texto al final de cada epoch para monitorear
la calidad cualitativa del modelo (detecta degeneracion, repeticion, etc.).*
- `evaluate` (line 2651) `def evaluate(self, dataloader)`
- `__init__` (line 2701) `def __init__(self, config)`
- `compute_delta` (line 2709) `def compute_delta(self, model)`
- `compute_alpha` (line 2716) `def compute_alpha(self, delta)`
- `update_grad_buffer` (line 2721) `def update_grad_buffer(self, model)` - *Captura gradientes de forma segura, ignorando tensores corruptos.*
- `compute_t_eff` (line 2747) `def compute_t_eff(self, lr)` - *T_eff = lr/2 * Var(gradiente). Temperatura termodinamica efectiva.*
- `compute_kappa` (line 2755) `def compute_kappa(self, model, dataloader, n_batches)` - *κ = λ_max / λ_min de la covarianza del gradiente.
Parámetro de orden para cristalización (κ≈1 = cristal).
Nota: requiere pasadas backward adicionales. Se ejecuta con protección
para no corromper el estado AMP del trainer principal.*
- `compute_berry_phase` (line 2813) `def compute_berry_phase(self, model)` - *Fase de Berry de los kernels espectrales imaginarios.
Surge de los parametros ki_w, ki_x, ki_y, ki_z de QuaternionSpectralLayer.
|berry|>pi/2 con winding!=0 indica estructura topologica.*
- `compute_lc` (line 2826) `def compute_lc(self, model)` - *Complejidad local: 1 - similitud coseno promedio entre filas de pesos.*
- `compute_sp` (line 2840) `def compute_sp(self, model)` - *Superposicion: correlacion inter-fila promedio (entrelazamiento de features).*
- `classify_phase` (line 2856) `def classify_phase(self, delta, kappa, berry)` - *Clasificacion de fase segun Book.md:

discrete_crystal:       delta<0.05, kappa<1.5
topological_insulator:  |berry|>pi/2, winding!=0
cold_glass:             kappa>>1, delta>0.3
functional_glass:       intermedio (lo mas comun en LM)*
- `compute_all` (line 2875) `def compute_all(self, model, lr, dataloader, compute_kappa)` - *Calcula todas las metricas.
compute_kappa=True hace pasadas backward adicionales (caro, usar cada N epochs).*
- `format_log` (line 2900) `def format_log(self, m)`
- `__init__` (line 2936) `def __init__(self, config, logger)`
- `_measure_ratio` (line 2940) `def _measure_ratio(self, ratio, sample_batch)` - *Mide la coherencia espectral para un ratio dado.
Retorna: varianza del gradiente (menor = mas coherente = mejor).*
- `optimize` (line 2969) `def optimize(self, dataloader)` - *Retorna el mejor ratio de inicializacion de kernels espectrales.*
- `__init__` (line 3009) `def __init__(self, config, logger)`
- `prospect` (line 3013) `def prospect(self, candidates, train_dataset, prospect_steps)` - *Retorna el mejor batch size segun delta y T_eff.*
- `__init__` (line 3092) `def __init__(self, config, logger)`
- `mine` (line 3096) `def mine(self, seed_start, n_seeds, train_dataset, prospect_steps)` - *Retorna la semilla con la mejor trayectoria de delta.*
- `__init__` (line 3178) `def __init__(self, trainer, t0, cooling_rate, stagnation_patience)`
- `refine` (line 3187) `def refine(self, train_dl, val_dl, refine_epochs)` - *Ejecuta refine_epochs epocas de recocido simulado.
Retorna el historial de refinamiento.*
- `__init__` (line 3330) `def __init__(self, config, train_tokens, val_tokens, tokenizer, logger, curriculum_tiers, progressive_seq)`
- `_build_dataloader` (line 3343) `def _build_dataloader(self, tokens, seq_len, batch_size, shuffle, tag)`
- `_build_phases` (line 3357) `def _build_phases(self)`
- `run` (line 3366) `def run(self, run_prospect, refine_epochs, resume, prospect_steps, probe_seeds, seed_start)`
- `__init__` (line 3468) `def __init__(self, config, train_dataset, val_dataset, tokenizer, logger)`
- `_make_dataloaders` (line 3478) `def _make_dataloaders(self, batch_size)`
- `run` (line 3490) `def run(self, run_prospect, refine_epochs, resume, prospect_steps, probe_seeds, seed_start)` - *Ejecuta el pipeline completo.
Retorna el trainer con el modelo entrenado.*
- `ckpt_fn` (line 1031) `def ckpt_fn(x_in)`

#### `train.py`
**Path:** `topogpt3/train.py`
**File Doc:** *TopoGPT3: Grassmannian / Berry-Holonomy extension of TopoGPT2  Author: Gris Iscomeback License: GPL v3  Lo nuevo respecto a model.py --------------------------------- 1. Espacio base: Grassmanniana Gr(r, N) sobre el tensor de kernels espectrales K(theta) en C^{N_f x N_c}. El estado geometrico vive en U_r(theta) en St(r,N)/U(r), con r elegido dinamicamente por el "elbow" del espectro singular de K. 2. Fisher gap funcional:    Delta_F(theta) = lambda_r(Sigma_F) - lambda_{r+1}(Sigma_F) estimado por covarianza empirica de gradientes (mini-batch) o por scores. 3. Conexion de Berry discreta:  A_n = i * U_n^dagger (U_{n+1} - U_n) y holonomia acumulada      U_Gamma = P prod_n exp(-i A_n)  en U(r). 4. Distancia de conjugacion en SU(2) (cuaternionico, r=1 efectivo): d_conj(U1, U2) = min_{g in SU(2)} || U1 - g U2 g^{-1} ||_F 5. Winding W como proxy heuristico barato (rol secundario). 6. Curriculum por dataset, de mas simple a mas dificil: Tier 1: CodeAlpaca               (instrucciones cortas) Tier 2: Code-Feedback-Filtered   (chat / explicacion paso a paso) Tier 3: Magicoder-Evol-Instruct-110K (problemas complejos) Tier 4: Tiny-The-Stack           (codigo real multilenguaje) Cada tier mantiene splits train / val / holdout *disjuntos*. El conjunto HOLDOUT nunca se ve durante entrenamiento; se usa solo para medir generalizacion verdadera al final de cada tier y al final del pipeline.  Diseno*

**Classes:**
- `TopoGPT3Config` (line 87) `class TopoGPT3Config` - *Configuracion del pipeline TopoGPT3 (Grassmanniana + curriculum).*
- `GrassmannianTracker` (line 209) `class GrassmannianTracker` - *Observables geometricos sobre la trayectoria SGD.

En cada snapshot:
  - Apila los kernels espectrales (kr_*, ki_*) del modelo en
    K(theta) en C^{N_f x N_c}.
  - SVD truncada -> U_r(theta) en St(r,N).
  - Rango r dinamico por elbow de los valores singulares.
  - Gap funcional Delta_F estimado por covarianza de gradientes
    muestrales (proxy de la matriz de Fisher).
  - Conexion de Berry discreta entre snapshots consecutivos:
       A_n = i * U_n^dagger (U_{n+1} - U_n)
    Holonomia acumulada U_Gamma = P prod_n exp(-i A_n) en U(r).
  - Distancia de conjugacion en SU(2) (r=1 efectivo cuaternionico).
  - Winding W como proxy barato.

Todos los calculos viven en CPU/float32 para no contaminar AMP.*
- `EfficiencyMetrics` (line 597) `class EfficiencyMetrics` - *Mide y calcula los tres ratios pedidos:

  perf_per_param  =  (1 / val_ppl) / params_M
  perf_per_FLOP   =  tokens_per_sec / FLOPs_per_sec_aprox
  perf_per_BW     =  tokens_per_sec / bytes_moved_per_sec_aprox

FLOPs estimados con la heuristica de Kaplan/Hoffmann:
    FLOPs_forward_per_token ~= 2 * N_no_embed
    FLOPs_total_per_token  ~= 6 * N_no_embed       (forward + backward)
Bandwidth estimada como params_bytes leidos + activations_bytes movidas por step.
tokens_per_sec se cronometra empiricamente sobre el dataloader.*
- `CodeCurriculumLoader` (line 725) `class CodeCurriculumLoader` - *    Carga los 4 datasets, normaliza cada ejemplo a una unica cadena de texto,
    tokeniza con BPE y produce splits train / val / holdout disjuntos.

    Politica de normalizacion por dataset:
      - CodeAlpaca:           "### Instruction
{i}
### Input
{x}
### Response
{o}"
      - Code-Feedback:        concat de turnos: "<usr> ... </usr>
<asst> ... </asst>"
      - Magicoder-Evol:       "### Problem
{p}
### Solution
{s}"
      - Tiny-The-Stack:       texto crudo del archivo (truncado a 32k chars/file)

    Cache en disco: tokens_{tier}_{split}.bin (int32 memmap) + manifest .json.
    El HOLDOUT se separa con seed fija antes de tokenizar para garantizar
    que la misma muestra nunca aparezca en train o val entre corridas.
    *
- `BlockTokenDataset` (line 1012) `class BlockTokenDataset(Dataset)` - *Dataset autoregresivo sobre un stream de tokens.
Cada item es (x, y) con shape [seq_len].*
- `CheckpointStore` (line 1039) `class CheckpointStore` - *Persiste pesos del modelo + estado del trainer (sin AMP scaler para portabilidad).*
- `TopoGPT3Trainer` (line 1115) `class TopoGPT3Trainer` - *Orquesta el curriculum sobre los 4 tiers.

Pipeline por tier:
  1. Abre memmap de tokens (train/val/holdout).
  2. Construye DataLoaders con seq_len(tier).
  3. Entrena TIER_EPOCHS[tier] epocas con AMP + grad accum.
  4. Cada GRASS_TRACK_EVERY steps: snapshot Grassmanniano.
  5. Al final de cada epoca: eval en VAL.
  6. Al final del tier: eval en HOLDOUT (datos nunca vistos).
  7. Checkpoint y avanza al siguiente tier.

Al final del pipeline: eval en HOLDOUT *combinado* de los 4 tiers.*

**Methods:**
- `_gauss_complex_contract` (line 544) `def _gauss_complex_contract(self, W, X)` - *Sustituye QuaternionSpectralLayer._contract usando el truco de Gauss.

Para (Wr + i Wi)(Xr + i Xi) la version naive requiere 4 productos reales:
    Yr = Wr Xr - Wi Xi
    Yi = Wr Xi + Wi Xr
Gauss (Karatsuba) baja a 3 productos reales:
    m1 = Wr * Xr
    m2 = Wi * Xi
    m3 = (Wr + Wi) * (Xr + Xi)
    Yr = m1 - m2
    Yi = m3 - m1 - m2

Importante (AMP): el _contract original opera sobre complex64 y PyTorch no
autocastea operaciones complejas; el resultado es complex64. Si dejamos que
autocast convierta nuestros einsums reales a fp16, la dtype de salida cambia
y rompe el scatter_add_ corriente abajo en QuaternionTorusBrain. Por eso
desactivamos autocast aqui y forzamos fp32 para preservar la semantica.*
- `apply_gauss_patch` (line 580) `def apply_gauss_patch(logger)` - *Activa la version Gauss de _contract en QuaternionSpectralLayer.
Idempotente: solo parchea una vez por proceso.*
- `parse_args` (line 1728) `def parse_args()`
- `main` (line 1764) `def main()`
- `build_topogpt2_config` (line 182) `def build_topogpt2_config(self, max_seq_len, attn_window)`
- `__init__` (line 229) `def __init__(self, config, logger)`
- `_stack_spectral_kernels` (line 243) `def _stack_spectral_kernels(model)` - *Devuelve K(theta) en C^{N_f x N_c}:
  - filas = frecuencias planas (todos los modos espaciales de todos los kernels)
  - columnas = canales (in_q * out_q por componente cuaternionico, sumados)*
- `_elbow_rank` (line 280) `def _elbow_rank(self, sigmas)` - *Punto donde el valor singular cae por debajo de elbow_ratio * sigma_max.*
- `_dominant_subspace` (line 289) `def _dominant_subspace(self, K)` - *SVD compacta y truncada.
Devuelve (U_r, sigmas, r) con U_r en C^{N_f x r} ortonormal.*
- `_flatten_grads` (line 307) `def _flatten_grads(model, max_per_tensor)` - *Concatena un sub-sample de gradientes para mantener costo acotado.*
- `estimate_fisher_gap` (line 325) `def estimate_fisher_gap(self, model, dataloader, vocab_size, r_target)` - *Sigma_F ~= (1/M) sum_m g_m g_m^T  (covarianza muestral de gradientes).
Delta_F = lambda_{r_eff} - lambda_{r_eff+1}, donde r_eff = min(r_target, M-2)
para no salir del rango efectivo del estimador con M gradientes.
Devuelve (gap, eigs_desc, r_eff).*
- `_project_unitary` (line 393) `def _project_unitary(M)` - *Proyeccion a U(r) por descomposicion polar (M ~= U H -> retorna U).*
- `update_holonomy` (line 398) `def update_holonomy(self, U_new)` - *Holonomia discreta:
    T_n = U_n^dagger U_{n+1}  en C^{r x r}  (transporte paralelo discreto)
    U_Gamma <- T_n * U_Gamma  (acumulado)
Tras cada paso, U_Gamma se proyecta a U(r) para evitar deriva numerica.*
- `conjugation_distance_su2` (line 424) `def conjugation_distance_su2(U1, U2)` - *Para U1, U2 en U(1)/U(2):  d_conj(U1, U2) = min_g || U1 - g U2 g^{-1} ||_F.
En U(1) coincide con |U1 - U2|.
En SU(2) se reduce a comparar |Tr(U1)| con |Tr(U2)| (clase de conjugacion).*
- `_accumulate_winding` (line 441) `def _accumulate_winding(self, U_new)` - *W += (1/2pi) * arg det <U_prev | U_new>  acumulado sobre la trayectoria.*
- `snapshot` (line 457) `def snapshot(self, model, step, dataloader, vocab_size)`
- `format_log` (line 512) `def format_log(self, snap)`
- `save` (line 534) `def save(self, path)`
- `__init__` (line 612) `def __init__(self, model, config, logger, gauss_enabled)`
- `_embed_params` (line 623) `def _embed_params(model)`
- `measure_throughput` (line 631) `def measure_throughput(self, dataloader, vocab_size)` - *Devuelve (tokens_por_segundo, segundos_por_step).*
- `estimate_flops_per_step` (line 663) `def estimate_flops_per_step(self, batch_size, seq_len)` - *Heuristica: 6 * N_no_embed * tokens (forward + backward).*
- `estimate_bytes_per_step` (line 668) `def estimate_bytes_per_step(self, batch_size, seq_len, dtype_bytes)` - *Bandwidth aproximada: lectura de pesos + activaciones por step.
Asume AMP fp16 (2 bytes); pesos fp32 (4 bytes) leidos una vez.*
- `compute` (line 676) `def compute(self, dataloader, vocab_size, val_loss, val_ppl, val_acc, batch_size, seq_len)`
- `format_log` (line 708) `def format_log(self, m)`
- `__init__` (line 741) `def __init__(self, config, tokenizer, logger)`
- `_format_codealpaca` (line 758) `def _format_codealpaca(ex)`
- `_format_code_feedback` (line 769) `def _format_code_feedback(ex)`
- `_format_magicoder` (line 789) `def _format_magicoder(ex)`
- `_format_tiny_stack` (line 797) `def _format_tiny_stack(ex)`
- `_get_formatter` (line 809) `def _get_formatter(cls, tier)`
- `_tier_paths` (line 831) `def _tier_paths(self, tier)`
- `_manifest_path` (line 837) `def _manifest_path(self, tier)`
- `_already_prepared` (line 840) `def _already_prepared(self, tier)` - *True solo si los 3 splits existen, son no-vacios y el manifest concuerda.*
- `_load_hf_with_fallback` (line 870) `def _load_hf_with_fallback(self, tier)` - *Carga el dataset HF; para tiny_the_stack prueba una cadena de fallbacks
publicos hasta que uno funcione.*
- `prepare_tier` (line 903) `def prepare_tier(self, tier_index, force)`
- `open_memmap` (line 998) `def open_memmap(self, tier, split)`
- `__init__` (line 1018) `def __init__(self, tokens, seq_len)`
- `__len__` (line 1023) `def __len__(self)`
- `__getitem__` (line 1026) `def __getitem__(self, idx)`
- `__init__` (line 1042) `def __init__(self, root, max_keep, logger)`
- `save` (line 1049) `def save(self, tag, model, optimizer, state)` - *Guarda checkpoint atomico en <root>/last/ sobreescribiendo el anterior.

El argumento `tag` se conserva por compatibilidad pero se ignora: solo
existe un checkpoint llamado `last` y los pesos en safetensors.*
- `load_latest` (line 1083) `def load_latest(self, model, optimizer)`
- `should_save` (line 1107) `def should_save(self, interval_min)`
- `__init__` (line 1131) `def __init__(self, config, start_tier, exploitgym_loader, topogpt3_loader, merged_config)`
- `prepare_all` (line 1203) `def prepare_all(self, force)`
- `_build_loaders` (line 1244) `def _build_loaders(self, tier_index)`
- `_cosine_lr` (line 1294) `def _cosine_lr(self, step, total_steps)`
- `_set_lr` (line 1301) `def _set_lr(self, lr)`
- `_train_one_tier` (line 1309) `def _train_one_tier(self, tier_index)`
- `_check_forgetting` (line 1527) `def _check_forgetting(self, completed_tier_index)` - *Evaluate holdout of ALL previous tiers to detect catastrophic forgetting.*
- `_open_memmaps` (line 1557) `def _open_memmaps(self, tier_index)` - *Open train/val/holdout memmaps for a given tier.*
- `_evaluate` (line 1585) `def _evaluate(self, dl)` - *Devuelve (avg_loss, perplexity, token_accuracy).*
- `_state_dict` (line 1623) `def _state_dict(self)`
- `run` (line 1636) `def run(self)`
- `_eval_combined_holdout` (line 1693) `def _eval_combined_holdout(self)`
- `flush` (line 934) `def flush(split)`

#### `vanilla_control.py`
**Path:** `topogpt3/vanilla_control.py`
**File Doc:** *Vanilla transformer control for TopoGPT3 Hodge-CM ablation. Matched size: D=384, L=8, H=8, SwiGLU FFN, RoPE, RMSNorm, tied embeddings. Target: ~33M to compare against TopoGPT3 37.4M (which includes TorusBrain+MoE+spectral). Same vocab (50257), full curriculum interface (codealpaca, code_feedback, magicoder_evol, tiny_the_stack) with train/val/holdout disjuntos — nada recortado. Usa tokens reales del tier cuando existen; sintetico solo con --allow-synthetic explícito.*

**Classes:**
- `VanillaConfig` (line 61) `class VanillaConfig`
- `RMSNorm` (line 72) `class RMSNorm(Module)`
- `RotaryEmbedding` (line 83) `class RotaryEmbedding(Module)`
- `CausalAttention` (line 113) `class CausalAttention(Module)`
- `SwiGLU` (line 134) `class SwiGLU(Module)`
- `VanillaLayer` (line 146) `class VanillaLayer(Module)`
- `VanillaTransformer` (line 161) `class VanillaTransformer(Module)` - *Tied-embedding decoder-only transformer. Same training interface as TopoGPT2.*

**Functions:**
- `resolve_device` (line 22) `def resolve_device(req)`
- `find_tier_tokens` (line 41) `def find_tier_tokens(tier)` - *Busca train real tokenizado del tier. Devuelve lista de ints o None.
Nada silencioso: si no hay, el llamador decide (error o sintetico explícito).*

**Methods:**
- `build_optimizer` (line 191) `def build_optimizer(model, cfg)`
- `__init__` (line 73) `def __init__(self, d, eps)`
- `forward` (line 78) `def forward(self, x)`
- `__init__` (line 84) `def __init__(self, d_head, max_seq_len, base)`
- `_build_cache` (line 90) `def _build_cache(self, seq_len)`
- `_rot_half` (line 98) `def _rot_half(x)`
- `forward` (line 102) `def forward(self, q, k)`
- `__init__` (line 114) `def __init__(self, cfg)`
- `forward` (line 124) `def forward(self, x)`
- `__init__` (line 135) `def __init__(self, d)`
- `forward` (line 142) `def forward(self, x)`
- `__init__` (line 147) `def __init__(self, cfg)`
- `forward` (line 155) `def forward(self, x)`
- `__init__` (line 164) `def __init__(self, cfg)`
- `_init` (line 176) `def _init(m)`
- `forward` (line 180) `def forward(self, ids)`
- `count_params` (line 186) `def count_params(self)`

#### `transfer_weights.py`
**Path:** `transfer_weights.py`
**File Doc:** *Transfer weights from a smaller TopoGPT model to a larger one.  Strategy: - Matching layers (same shape): direct copy - Embedding/lm_head (vocab same, dim larger): pad with zeros - Linear layers (in or out dim larger): copy sub-block, pad rest - QuaternionSpectralLayer kernels: copy matching freq bins, pad rest - New layers (no match): random init (default torch init)  This gives the larger model a head start from the smaller model's knowledge.  Usage: python transfer_weights.py --from checkpoints_topogpt3/last/model.safetensors \\ --to checkpoints_topoexploit/last/model.safetensors \\ --scale-from small --scale-to large*

**Functions:**
- `transfer_weights` (line 31) `def transfer_weights(state_small, model_large, logger)` - *Transfer weights from small model state_dict to large model.
Returns the new state_dict for the large model.*
- `_copy_with_padding` (line 135) `def _copy_with_padding(small, large)` - *Copy small tensor into the top-left of a larger tensor, rest stays as large's init.*
- `main` (line 147) `def main()`

### SH (5 files)

#### `install.sh`
**Path:** `install.sh`

*No symbols extracted*

#### `run_exploitgym.sh`
**Path:** `run_exploitgym.sh`
**File Doc:** *TopoExploit: 125M params trained on ExploitGym  Usage: ./run_exploitgym.sh prepare     -- clone repo + tokenize data ./run_exploitgym.sh transfer    -- transfer weights from small model ./run_exploitgym.sh train       -- train from scratch ./run_exploitgym.sh train-init  -- train with transferred weights ./run_exploitgym.sh eval        -- evaluate on holdout ./run_exploitgym.sh full        -- prepare + transfer + train + eval*

*No symbols extracted*

#### `run_exploitgym_v2.sh`
**Path:** `run_exploitgym_v2.sh`

**Functions:**
- `banner` (line 24)
- `train_v2` (line 32)
- `eval_v2` (line 45)
- `infer_v2` (line 54)

#### `run_hodge_cm_ablation.sh`
**Path:** `run_hodge_cm_ablation.sh`
**File Doc:** *Hodge-CM offline ablation runner (GPU si el build torch lo permite, si no CPU). No recorta nada: 12 kernels reales + probe funcional + control vanilla con curriculum completo (codealpaca/code_feedback/magicoder/tiny_the_stack, train/val/holdout disjuntos). Sin sintetico silencioso. Usage: ./run_hodge_cm_ablation.sh [--device auto|cuda|cpu]*

*No symbols extracted*

#### `run_merged.sh`
**Path:** `run_merged.sh`

**Functions:**
- `kill_gpu_processes` (line 33)
- `restore_gpu_processes` (line 63)
- `banner` (line 84)
- `prepare_data` (line 93)
- `train_merged` (line 107)
- `eval_merged` (line 119)
- `infer_merged` (line 128)
