Fixed-budget multi-agent inference benchmark harness for studying when split inference helps or hurts versus a strong single-agent baseline under local context ceilings, using local Ollama gemma3:1b, CPU-only pilots, topology diagnostics, verification-budget analysis, and reproducible experiment logging.
ai topology multi-agent-systems split-inference mixture-of-experts ai-evaluation llm-benchmark inference-allocation fixed-budget local-context experiment-harness
-
Updated
Mar 12, 2026 - Python