

Learning Layer Paper Reading Club - Week 30 - ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
This week's paper: ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Link: https://arxiv.org/abs/2605.03042 Code: wanshuiyin/Auto-claude-code-research-in-sleep
Abstract:
How well an LLM agent performs depends not just on the model weights but on the harness around them — what gets stored, retrieved, and put in front of the model. Over a long-horizon research workflow, the dangerous failure isn't a crash. It's a plausible unsupported success: an agent that runs for hours and produces claims whose evidence is incomplete, misreported, or silently inherited from its own framing.
ARIS (Auto-Research-in-sleep) is an open-source research harness that attacks this with cross-model adversarial collaboration. An executor model drives progress forward while a reviewer from a different model family critiques intermediate artifacts and demands revisions.
It has three layers. The execution layer holds 65+ reusable Markdown-defined skills, model integrations over MCP, a persistent research wiki, and deterministic figure generation. The orchestration layer runs five end-to-end workflows with adjustable effort and configurable reviewer routing. The assurance layer checks whether claims are actually supported by evidence — integrity verification, result-to-claim mapping, and claim auditing that cross-checks the manuscript against a claim ledger and the raw evidence — plus a five-pass scientific-editing pipeline, proof checks, and visual inspection of the rendered PDF. A prototype self-improvement loop records research traces and proposes harness changes, adopted only after reviewer approval.
Discussion Topics:
Why the harness, not the model, is often the binding constraint on agent performance
"Plausible unsupported success" as the central long-horizon failure mode
Whether cross-family adversarial review actually catches what same-family review misses
Claim ledgers and result-to-claim mapping as a check on hallucinated results
Markdown-defined skills and persistent wikis as agent memory
What a self-improving harness can and cannot safely change about itself
Relevant if you're building agents, evaluation harnesses, or anything that has to trust an LLM's own report of what it did. No preparation required — come read and discuss with us.
Mat will lead the discussion this week
What is the Learning Layer Labs Paper Reading Club?
An initiative from https://www.learninglayer.ai, a lab with the goal of reducing AI anxiety in the world.
What is the format?
Discussion based. Expect a low pressure environment to share insights and opinions with the group.
What are the group goals?
Stay on top of AI research and improve understanding of AI fundamentals + math.
Who is welcome?
Everyone! Try to put in at least some time on the paper and come prepared with questions or things you'd like to discuss, but it's ok to just show up!
Learning Layer Labs team:
Thomas Redfern ( The backbone who runs every paper reading )
Mat Allen (The new fella bring order and organization)
Devinder Sodhi (The guy you talk to for sponsoring)
The FOLD SF is hosting us this week!
About The Fold SF
Started in 2025 by long time SF community folks looking to convene, organize, and explore in new ways. A space to center different possibilities.
The Fold continues what was previously The Laundry in a new chapter.