Research

SmartInfer Research studies how specialised AI systems and quantitative methods can support consequential decisions and perform substantial work without claiming more than the evidence supports. Our work runs in three directions: AI agents under explicit verification; smaller models trained and evaluated for bounded domains; and measurement methods that state only what the data can support.

We publish experimental reports alongside research papers and technical foundations. Each experimental report states the question, the evidence used to evaluate it, what the result demonstrates, and its limits.

Research directions

Original SmartInfer work across three research directions. Lead reports state the question, finding, and strongest limitation; supporting papers develop the underlying methods and theory.

AI agents

State-changing work should be accepted through evidence independent of the agent’s own account.

Building agents that complete substantial work under explicit verification.

Verified Software Generation: The First Protozoa Experiments

Experimental report · Protozoa · July 2026 · two end-to-end systems

Question
Can generated software be built, accepted, and repaired through explicit, checkable verification gates?
Finding
Protozoa produced and verified two complete software packages. In the browser-game experiment, a human-observed defect became a regression condition that failed on the flawed artifact, passed after bounded repair, and left the existing test suite green.
Limit
Two early, human-directed experiments; not evidence of general autonomous software engineering.

Supporting papers

  • Untrusted Builders, Trusted Gates: A Verification Architecture for Agentic Software

    Conference paper · submitted

    Proposes that untrusted generators may produce patches and repairs, while acceptance remains controlled by independent, machine-checkable verification gates.

  • Position: Agent Verification Requires Typed Workflows and Bounded Commitments

    Conference paper · submitted

    Argues that agent verification should evaluate committed traces over typed workflow actions, with explicit gates separating open-ended proposals from state-changing commitments.

  • Verifying Structured Retrieval with Laplacian Coherence

    Conference paper · submitted

    Develops a label-free method for estimating retrieval confidence through graph-local coherence over evidence candidates.

Current work studies agents whose state-changing actions are accepted through independently checkable tests, properties, contracts, or formal methods.

AI models

Specialist capability becomes testable when the domain, teaching signal, and evaluation are bounded together.

Studying when smaller models can acquire useful specialist capability in bounded domains.

Teaching a Tiny Model a New Language

Experimental report · Velk · July 2026 · frozen 100-problem test

Question
Can a 494M-parameter model acquire a genuinely new bounded skill from self-generated, machine-verified data?
Finding
After three training rounds using self-generated, machine-verified traces, a 494M open-weight model improved from 10% to 23% on a frozen 100-problem test. The gain occurred in non-guessable task types; a full-manual-in-context baseline remained at 10%.
Limit
Single run with no seed replication; the direct-fine-tuning control remains unscored, and the preregistered 40%/20% target was not met.

Supporting papers

  • The Information Geometry and Sample Complexity of Grouped Preference Optimization

    Journal paper · submitted

    Analyzes when larger comparison groups improve listwise preference learning and when evaluation noise becomes the controlling factor in sample complexity.

Velk is a measuring instrument, not a product domain. Current work studies when bounded, machine-verifiable teaching signals can create useful specialist capability in small models.

Measurement

An estimate is useful for decisions only when the evidence supports acting on it.

Developing estimators, credibility gates, and experiment designs that state what the data can and cannot support.

From Model Fit to Marketing Decisions

Experimental report · Galileo Measurement Science 0.2 · August 2026 · 90 synthetic datasets

Question
Do fitted marketing response curves preserve decision value, rather than merely reproduce historical outcomes?
Finding
Across 90 synthetic datasets with known economic ground truth, Galileo 0.2 had the lowest median normalized decision regret among five evaluated methods. It did not lead every data regime, supporting the conclusion that no method should be treated as universally superior.
Limit
Synthetic benchmark; it does not establish real-customer lift, real-world estimator superiority, or causal incrementality in observational customer data.

The Confident Wrong Number: Identifiability in Marketing Mix Modeling

Research note · Marketing science · July 2026 · synthetic demonstration

Question
What happens when weakly identifying data produce precise-looking attribution?
Finding
On a synthetic brand where media truly contributed 15.1% of revenue, a standard regularized fit attributed 97.0% and marked the result credible. An identifiability-gated fit still over-attributed at 31.3%, but returned a NOT CREDIBLE verdict—illustrating why the credibility decision can matter more than a precise-looking number.
Limit
Synthetic ground-truth demonstration; no real-world accuracy or comparative-superiority claim.

Supporting papers

  • Exploration Is Bounded by Estimation: Fisher-Directed Budget Allocation under Hill Saturation

    Journal paper · in preparation

    Develops a method for allocating exploration budget under saturating channel-response curves by identifying where additional spend can most improve estimation.

Current work extends observational estimation toward designed experiments that identify which additional evidence would most reduce decision-relevant uncertainty.

Foundations & surveys

Selected technical surveys and earlier writing by SmartInfer founder Anjan Goswami, presented as background to SmartInfer’s current research.

Memory & agents

Search & retrieval

Reasoning & reinforcement learning

Enterprise AI architecture

Industry context

More technical writing: technical articles. Academic work: selected publications, DBLP bibliography, and patents.