Working papers, benchmarks, and experimental notes. The thread connecting them: evidence-based AI systems — retrieval, verification, synthesis, and measurement for agentic pipelines.

Evidence-Weighted Routing and Error Measurement for Code Localization and Agentic Repair: A Three-Phase Study

Attention and embedding retrieval are good at finding candidate regions, but the evidence needed to justify a requirement-satisfaction claim is relational, so the final support belongs in a graph or evidence subgraph. The paper covers hierarchical evidence routing, traceability drift as evidence-weighted innovation, and a frozen Codex augmentation replay.

Benchmarks & Experiments

Conceptual Framework

Reproducibility