# eval: harness

*Community 1 | 7 files | cohesion 0.64*

## Definition

This community groups 7 file(s) rooted at `eval` with dominant language py (cohesion 0.64). Central symbols: `ModelLoader`, `SandboxConfig`, `__init__`, `_blocked_dunder_access`, `_build_worker_src`, `_is_env_truthy`, `_make_hrm`, `_make_standard`. Core file: `eval/harness.py` (12 symbols). Documented purpose: Harness for evaluating TopoGPT3 on HumanEval (164 problems).  Faithful to the official HumanEval protocol: for each problem we feed the model the function signa.

## Files

| File | Language | Layer | Symbols | Doc |
|------|----------|-------|---------|-----|
| `eval/harness.py` | py | utility | 12 | yes |
| `eval/integration_smoke.py` | py | utility | 1 | yes |
| `eval/noise_sweep.py` | py | utility | 4 | yes |
| `eval/samplers.py` | py | utility | 7 | yes |
| `eval/sandbox.py` | py | utility | 9 | yes |
| `eval/sandbox_smoke.py` | py | utility | 1 | yes |
| `eval/temp_sweep.py` | py | utility | 5 | yes |

## Key Symbols

- `load_humaneval` (function, `eval/harness.py:59`) `def load_humaneval(cache_dir)`
- `build_prompt` (function, `eval/harness.py:75`) `def build_prompt(problem)` - Return the exact prompt text fed to the model.
- `extract_candidate` (function, `eval/harness.py:100`) `def extract_candidate(prompt, completion)` - Combine prompt + completion into a single Python source string.
- `run_one_test` (function, `eval/harness.py:150`) `def run_one_test(problem, candidate_src, timeout)` - Execute the candidate against the hidden test.
- `run_one_test_sandboxed` (function, `eval/harness.py:172`) `def run_one_test_sandboxed(problem, candidate_src, timeout, sandbox_cfg)` - Sandboxed variant of `run_one_test`. Runs the candidate in a
- `make_sampler` (function, `eval/harness.py:195`) `def make_sampler(mode, settings_kwargs)` - Backwards-compatible shim. The real implementation lives in
- `completion_for_problem` (function, `eval/harness.py:204`) `def completion_for_problem(sampler, prompt)` - Run a single completion and return (raw_output_text, metrics_dict).
- `ModelLoader` (class, `eval/harness.py:217`) `class ModelLoader` - Build the model and tokenizer once, run many generations.
- `__init__` (method, `eval/harness.py:220`) `def __init__(self, ckpt_dir, ckpt_name, device)`
- `generate` (method, `eval/harness.py:246`) `def generate(self, prompt, max_new_tokens, temperature, top_k, repetition_penalt`
- `evaluate_problem` (method, `eval/harness.py:272`) `def evaluate_problem(problem, loader, args, sample_idx)`
- `main` (method, `eval/harness.py:315`) `def main()`
- `main` (function, `eval/integration_smoke.py:18`) `def main()`
- `inject_noise` (function, `eval/noise_sweep.py:46`) `def inject_noise(model, sigma, seed)` - Anade N(0, sigma) a TODOS los kernels espectrales (kr_*, ki_*).
- `load_model` (function, `eval/noise_sweep.py:74`) `def load_model(ckpt_dir, ckpt_name, device)` - Reconstruye TopoGPT2 alineado con el checkpoint, sin acceso a
- `generate_one` (function, `eval/noise_sweep.py:99`) `def generate_one(model, tok, prompt, max_new_tokens, device)`
- `main` (function, `eval/noise_sweep.py:117`) `def main()`
- `register_sampler` (function, `eval/samplers.py:36`) `def register_sampler(name)` - Decorator. Register a factory under `name`. If `enabled_env` is set,
- `deco` (function, `eval/samplers.py:42`) `def deco(fn)`
- `_is_env_truthy` (function, `eval/samplers.py:55`) `def _is_env_truthy(name)`
- `_make_standard` (function, `eval/samplers.py:64`) `def _make_standard(settings_kwargs)`
- `_make_hrm` (function, `eval/samplers.py:69`) `def _make_hrm(settings_kwargs)`
- `list_samplers` (function, `eval/samplers.py:86`) `def list_samplers()`
- `build_sampler` (function, `eval/samplers.py:90`) `def build_sampler(mode, settings_kwargs)` - Construct a sampler. Drop-in replacement for the old
- `SandboxConfig` (class, `eval/sandbox.py:53`) `class SandboxConfig` - One knob per defence layer. Defaults match HumanEval-style eval.
- `_names_imported` (method, `eval/sandbox.py:100`) `def _names_imported(tree)` - Return the set of top-level names brought into scope by imports.
- `_blocked_dunder_access` (method, `eval/sandbox.py:114`) `def _blocked_dunder_access(tree, blocked)` - Find Attribute nodes whose attr is in `blocked`. Returns attr names found.
- `_max_depth` (method, `eval/sandbox.py:123`) `def _max_depth(tree)` - Compute max nesting depth of the AST. Catches obfuscated huge trees.
- `d` (method, `eval/sandbox.py:125`) `def d(node, cur)`
- `check_safety` (method, `eval/sandbox.py:133`) `def check_safety(source, cfg)` - Return (ok, reason). `reason` is "" when ok, else a human-readable

## Internal vs External Edges

- Internal resolved imports (EXTRACTED): 8
- Cross-boundary resolved imports (EXTRACTED): 4

## Connections

- [EXTRACTED] depends_on community 1 <-> 2 (strength 0.9): Extracted import edge crosses communities: eval/harness.py imports topogpt3/__init__.py.
- [EXTRACTED] depends_on community 1 <-> 0 (strength 0.9): Extracted import edge crosses communities: eval/noise_sweep.py imports topogpt3/model.py.
- [INFERRED] bridges community 2 <-> 1 (strength 0.5): Inferred cross-community bridge: eval/governor.py reaches eval/sandbox_smoke.py in 5 hops.
- [INFERRED] bridges community 0 <-> 1 (strength 0.5): Inferred cross-community bridge: eval/hodge_cm_ablation.py reaches eval/sandbox_smoke.py in 5 hops.
- [INFERRED] bridges community 1 <-> 0 (strength 0.5): Inferred cross-community bridge: eval/integration_smoke.py reaches topogpt3/hodge_cm.py in 5 hops.
- [INFERRED] shares_context community 1 <-> 3 (strength 0.5): Inferred shared context (language py and layer utility) with no import path between community 1 (eval: harness) and community 3 (topogpt3: inference_hrm).
- [INFERRED] shares_context community 1 <-> 4 (strength 0.5): Inferred shared context (language py) with no import path between community 1 (eval: harness) and community 4 (topogpt3: jlens).
- [INFERRED] shares_context community 1 <-> 5 (strength 0.5): Inferred shared context (language py and layer utility) with no import path between community 1 (eval: harness) and community 5 (orphans).
- [INFERRED] bridges community 1 <-> 0 (strength 0.4): Inferred cross-community bridge: eval/sandbox_smoke.py reaches topogpt3/hodge_cm.py in 6 hops.

## Risks

- [taint high] `eval/harness.py` -> `eval/harness.py` via `subprocess` (0 hops)
- [taint high] `eval/harness.py` -> `eval/samplers.py` via `subprocess` (1 hops)
- [taint high] `eval/harness.py` -> `eval/sandbox.py` via `subprocess` (1 hops)
- [taint high] `eval/harness.py` -> `topogpt3/__init__.py` via `subprocess` (1 hops)
- [taint high] `eval/harness.py` -> `topogpt3/merged_config.py` via `subprocess` (2 hops)

## Open Questions

- Is the dangerous import `subprocess` in `eval/harness.py` still required, or can it be isolated?
- What would break if the most connected file in eval: harness changed?
- Should eval: harness be split, given cohesion 0.64?

## Sources

- `eval/harness.py`
- `eval/integration_smoke.py`
- `eval/noise_sweep.py`
- `eval/samplers.py`
- `eval/sandbox.py`
- `eval/sandbox_smoke.py`
- `eval/temp_sweep.py`
