Research
SmartInfer Research studies how specialised AI systems and quantitative methods can
support consequential decisions and perform substantial work without claiming more
than the evidence supports. Our work runs in three directions: AI agents under
explicit verification; smaller models trained and evaluated for bounded domains; and
measurement methods that state only what the data can support.
We publish experimental reports alongside research papers and technical foundations.
Each experimental report states the question, the evidence used to evaluate it, what
the result demonstrates, and its limits.
Research directions
Original SmartInfer work across three research directions. Lead reports state
the question, finding, and strongest limitation; supporting papers develop the
underlying methods and theory.
AI agents
State-changing work should be accepted through
evidence independent of the agent’s own account.
Building agents that complete substantial work under
explicit verification.
Experimental report · Protozoa · July 2026 · two end-to-end systems
- Question
- Can generated software be built, accepted, and repaired through explicit, checkable verification gates?
- Finding
- Protozoa produced and verified two complete software packages. In the browser-game experiment, a human-observed defect became a regression condition that failed on the flawed artifact, passed after bounded repair, and left the existing test suite green.
- Limit
- Two early, human-directed experiments; not evidence of general autonomous software engineering.
Read report →
Supporting papers
-
Untrusted Builders, Trusted Gates: A Verification Architecture for Agentic Software
Conference paper · submitted
Proposes that untrusted generators may produce patches and repairs, while acceptance remains controlled by independent, machine-checkable verification gates.
-
Position: Agent Verification Requires Typed Workflows and Bounded Commitments
Conference paper · submitted
Argues that agent verification should evaluate committed traces over typed workflow actions, with explicit gates separating open-ended proposals from state-changing commitments.
-
Verifying Structured Retrieval with Laplacian Coherence
Conference paper · submitted
Develops a label-free method for estimating retrieval confidence through graph-local coherence over evidence candidates.
Current work studies agents whose state-changing actions are accepted through independently checkable tests, properties, contracts, or formal methods.
AI models
Specialist capability becomes testable when the
domain, teaching signal, and evaluation are bounded together.
Studying when smaller models can acquire useful specialist
capability in bounded domains.
Experimental report · Velk · July 2026 · frozen 100-problem test
- Question
- Can a 494M-parameter model acquire a genuinely new bounded skill from self-generated, machine-verified data?
- Finding
- After three training rounds using self-generated, machine-verified traces, a 494M open-weight model improved from 10% to 23% on a frozen 100-problem test. The gain occurred in non-guessable task types; a full-manual-in-context baseline remained at 10%.
- Limit
- Single run with no seed replication; the direct-fine-tuning control remains unscored, and the preregistered 40%/20% target was not met.
Read report →
Supporting papers
Velk is a measuring instrument, not a product domain. Current work studies when bounded, machine-verifiable teaching signals can create useful specialist capability in small models.
Measurement
An estimate is useful for decisions only when the
evidence supports acting on it.
Developing estimators, credibility gates, and experiment
designs that state what the data can and cannot support.
Experimental report · Galileo Measurement Science 0.2 · August 2026 · 90 synthetic datasets
- Question
- Do fitted marketing response curves preserve decision value, rather than merely reproduce historical outcomes?
- Finding
- Across 90 synthetic datasets with known economic ground truth, Galileo 0.2 had the lowest median normalized decision regret among five evaluated methods. It did not lead every data regime, supporting the conclusion that no method should be treated as universally superior.
- Limit
- Synthetic benchmark; it does not establish real-customer lift, real-world estimator superiority, or causal incrementality in observational customer data.
Read report →
Research note · Marketing science · July 2026 · synthetic demonstration
- Question
- What happens when weakly identifying data produce precise-looking attribution?
- Finding
- On a synthetic brand where media truly contributed 15.1% of revenue, a standard regularized fit attributed 97.0% and marked the result credible. An identifiability-gated fit still over-attributed at 31.3%, but returned a NOT CREDIBLE verdict—illustrating why the credibility decision can matter more than a precise-looking number.
- Limit
- Synthetic ground-truth demonstration; no real-world accuracy or comparative-superiority claim.
Read note →
Supporting papers
Current work extends observational estimation toward designed experiments that identify which additional evidence would most reduce decision-relevant uncertainty.
Foundations & surveys
Selected technical surveys and earlier writing by SmartInfer founder Anjan
Goswami, presented as background to SmartInfer’s current research.
Memory & agents
Search & retrieval
Reasoning & reinforcement learning
Enterprise AI architecture
Industry context
More technical writing: technical articles. Academic work: selected publications, DBLP bibliography, and patents.