<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Evaluation on 00n.ai | Architecture &amp; Research</title><link>https://00n.ai/tags/evaluation/</link><description>Recent content in Evaluation on 00n.ai | Architecture &amp; Research</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 26 Jun 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://00n.ai/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>Coding Knowledge Graph Agent Benchmark</title><link>https://00n.ai/research/coding-kg-agent-benchmark/</link><pubDate>Fri, 26 Jun 2026 00:00:00 +0000</pubDate><guid>https://00n.ai/research/coding-kg-agent-benchmark/</guid><description>Does a code knowledge graph with multi-agent navigation help small models write correct code? 4 models × 4 conditions × 9 tasks × 5 runs.</description></item><item><title>Knowledge Graph Agent Benchmark</title><link>https://00n.ai/research/knowledge-graph-agent-benchmark/</link><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><guid>https://00n.ai/research/knowledge-graph-agent-benchmark/</guid><description>Can a reasoning graph close the gap between small (7-8B) and frontier LLMs on statutory reasoning? Multi-agent, saturation, and synthesis scaffolding experiments.</description></item><item><title>Related Research: Multi-Agent Framework Benchmark Validation</title><link>https://00n.ai/research/related-research/</link><pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate><guid>https://00n.ai/research/related-research/</guid><description>Mapping our multi-agent benchmark findings to verified papers from arXiv, ACL, NeurIPS, ICML, and EMNLP. 17 papers evaluated for support, contradiction, and nuance.</description></item><item><title>Multi-Agent Framework Benchmark</title><link>https://00n.ai/research/multi-agent-framework-benchmark/</link><pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate><guid>https://00n.ai/research/multi-agent-framework-benchmark/</guid><description>Comparison data for vanilla, search, reflection, and multi-agent answer pipelines on freshness questions.</description></item></channel></rss>