# topogpt3: model

*Community 0 | 12 files | cohesion 0.64*

## Definition

This community groups 12 file(s) rooted at `topogpt3` with dominant language py (cohesion 0.64). Central symbols: `BPETokenizer`, `BlockTokenDataset`, `CMPhaseQuantSTE`, `CheckpointManager`, `CheckpointStore`, `CodeCurriculumLoader`, `CurriculumDataset`, `CurriculumTrainer`. Core file: `topogpt3/model.py` (187 symbols). Documented purpose: Diagnostico estatico de un checkpoint TopoGPT3 congelado.  Calcula sobre los pesos espectrales congelados (sin reentrenar):  kappa_F   = sigma_max / sigma_min  .

## Files

| File | Language | Layer | Symbols | Doc |
|------|----------|-------|---------|-----|
| `eval/diag_static.py` | py | utility | 5 | yes |
| `eval/hodge_cm_ablation.py` | py | utility | 6 | yes |
| `infer_exploitgym.py` | py | utility | 11 | yes |
| `synthetic_dataset.py` | py | data_access | 35 | yes |
| `topogpt3/ewc.py` | py | utility | 11 | yes |
| `topogpt3/exploitgym_config.py` | py | infrastructure | 3 | yes |
| `topogpt3/exploitgym_loader.py` | py | utility | 19 | yes |
| `topogpt3/hodge_cm.py` | py | utility | 15 | yes |
| `topogpt3/merged_config.py` | py | infrastructure | 6 | yes |
| `topogpt3/model.py` | py | business_logic | 187 | yes |
| `topogpt3/train.py` | py | utility | 64 | yes |
| `transfer_weights.py` | py | utility | 3 | yes |

## Key Symbols

- `phase_discretization` (function, `eval/diag_static.py:49`) `def phase_discretization(K, n_samples, seed)` - Muestrea n_samples overlaps aleatorios <u_i \| u_j> sobre los vectores
- `synthetic_winding` (function, `eval/diag_static.py:95`) `def synthetic_winding(K, n_windows, window_size)` - Como el checkpoint es estatico, no hay trayectoria temporal.
- `static_kappa` (function, `eval/diag_static.py:144`) `def static_kappa(K)`
- `context_length_diagnostic` (function, `eval/diag_static.py:171`) `def context_length_diagnostic(model, tracker, device, lengths)`
- `main` (function, `eval/diag_static.py:248`) `def main()`
- `resolve_device` (function, `eval/hodge_cm_ablation.py:28`) `def resolve_device(req)`
- `mat_stats` (function, `eval/hodge_cm_ablation.py:48`) `def mat_stats(kr, ki, r)`
- `heat_smooth` (function, `eval/hodge_cm_ablation.py:64`) `def heat_smooth(kr, ki, t, device)`
- `sm` (function, `eval/hodge_cm_ablation.py:70`) `def sm(x)`
- `main` (function, `eval/hodge_cm_ablation.py:74`) `def main()`
- `mean` (function, `eval/hodge_cm_ablation.py:141`) `def mean(k, m)`
- `load_task_ids` (function, `infer_exploitgym.py:41`) `def load_task_ids()`
- `load_task` (function, `infer_exploitgym.py:49`) `def load_task(task_id)` - Return a dict with keys: task_id, family, name, description, patch, pov.
- `build_prompt` (function, `infer_exploitgym.py:122`) `def build_prompt(task_info, tier)`
- `load_model` (function, `infer_exploitgym.py:136`) `def load_model(ckpt_dir, ckpt_name, device)` - Load tokenizer, build TopoGPT-2, apply Gauss patch, load weights.
- `generate` (function, `infer_exploitgym.py:177`) `def generate(model, tokenizer, prompt)` - Autoregressive generation with streaming + n-gram repetition blocking.
- `sample` (function, `infer_exploitgym.py:205`) `def sample(logits, seen_ids)`
- `run_single_prompt` (function, `infer_exploitgym.py:257`) `def run_single_prompt(args, model, tokenizer)`
- `run_eval_holdout` (function, `infer_exploitgym.py:274`) `def run_eval_holdout(args, model, tokenizer)`
- `run_interactive` (function, `infer_exploitgym.py:331`) `def run_interactive(args, model, tokenizer)`
- `parse_args` (function, `infer_exploitgym.py:380`) `def parse_args()`
- `main` (function, `infer_exploitgym.py:420`) `def main()`
- `LLMBackend` (class, `synthetic_dataset.py:61`) `class LLMBackend` - Abstract LLM backend. Subclass for each provider.
- `generate` (method, `synthetic_dataset.py:64`) `def generate(self, prompt)`
- `name` (method, `synthetic_dataset.py:67`) `def name(self)`
- `GroqBackend` (class, `synthetic_dataset.py:71`) `class GroqBackend(LLMBackend)` - Groq API backend using requests.
- `__init__` (method, `synthetic_dataset.py:78`) `def __init__(self, model, api_key, max_tokens, temperature, timeout)`
- `name` (method, `synthetic_dataset.py:95`) `def name(self)`
- `generate` (method, `synthetic_dataset.py:98`) `def generate(self, prompt)`
- `OpenRouterBackend` (class, `synthetic_dataset.py:121`) `class OpenRouterBackend(LLMBackend)` - OpenRouter unified API backend.

## Internal vs External Edges

- Internal resolved imports (EXTRACTED): 18
- Cross-boundary resolved imports (EXTRACTED): 12

## Connections

- [EXTRACTED] depends_on community 0 <-> 2 (strength 0.9): Extracted import edge crosses communities: eval/diag_static.py imports topogpt3/__init__.py.
- [EXTRACTED] depends_on community 1 <-> 0 (strength 0.9): Extracted import edge crosses communities: eval/noise_sweep.py imports topogpt3/model.py.
- [EXTRACTED] depends_on community 4 <-> 0 (strength 0.9): Extracted import edge crosses communities: tests/test_lens_model.py imports topogpt3/model.py.
- [EXTRACTED] depends_on community 3 <-> 0 (strength 0.9): Extracted import edge crosses communities: topogpt3/__main__.py imports topogpt3/train.py.
- [INFERRED] bridges community 2 <-> 0 (strength 0.5): Inferred cross-community bridge: eval/governor.py reaches topogpt3/hodge_cm.py in 5 hops.
- [INFERRED] bridges community 0 <-> 1 (strength 0.5): Inferred cross-community bridge: eval/hodge_cm_ablation.py reaches eval/sandbox_smoke.py in 5 hops.
- [INFERRED] bridges community 1 <-> 0 (strength 0.5): Inferred cross-community bridge: eval/integration_smoke.py reaches topogpt3/hodge_cm.py in 5 hops.
- [INFERRED] shares_context community 0 <-> 5 (strength 0.5): Inferred shared context (language py and layer utility) with no import path between community 0 (topogpt3: model) and community 5 (orphans).
- [INFERRED] bridges community 1 <-> 0 (strength 0.4): Inferred cross-community bridge: eval/sandbox_smoke.py reaches topogpt3/hodge_cm.py in 6 hops.

## Risks

- [taint critical] `eval/governor_smoke.py` -> `topogpt3/merged_config.py` via `eval` (2 hops)
- [taint critical] `eval/governor_smoke.py` -> `topogpt3/train.py` via `eval` (2 hops)
- [taint critical] `eval/governor_smoke.py` -> `topogpt3/model.py` via `eval` (2 hops)
- [taint critical] `eval/governor_smoke.py` -> `topogpt3/exploitgym_loader.py` via `eval` (3 hops)
- [taint critical] `eval/governor_smoke.py` -> `topogpt3/ewc.py` via `eval` (3 hops)
- [taint critical] `eval/governor_smoke.py` -> `topogpt3/exploitgym_config.py` via `eval` (3 hops)
- [taint critical] `eval/governor_smoke.py` -> `synthetic_dataset.py` via `eval` (3 hops)
- [taint high] `eval/harness.py` -> `topogpt3/merged_config.py` via `subprocess` (2 hops)

## Open Questions

- What would break if the most connected file in topogpt3: model changed?
- Should topogpt3: model be split, given cohesion 0.64?

## Sources

- `eval/diag_static.py`
- `eval/hodge_cm_ablation.py`
- `infer_exploitgym.py`
- `synthetic_dataset.py`
- `topogpt3/ewc.py`
- `topogpt3/exploitgym_config.py`
- `topogpt3/exploitgym_loader.py`
- `topogpt3/hodge_cm.py`
- `topogpt3/merged_config.py`
- `topogpt3/model.py`
- `topogpt3/train.py`
- `transfer_weights.py`
