Skip to content
Tallyn

How the AI works

Model-agnostic, and honest about it

Tallyn does not exist without modern language models, and it is not tied to any one lab. Here is the routing, the escalation rule, the fallback chain, and the check that decides what you are allowed to see.

Why a model at all

Because the judgement is the job

A rules engine can find an image with no alt attribute. It cannot tell you whether that alt should be empty because the image is decorative, or should read like a sentence because it carries the meaning of the paragraph next to it. It cannot tell a clickable div that traps a keyboard from one that is harmless, because that depends on what the div does.

The second half is harder. Explaining, in four sentences, who a barrier blocks and why, and then writing the fix for this team's own markup, is reasoning and writing. That is the part customers pay for and it is the part only a language model does.

So the model is the engine. Everything else in this product is plumbing around it: fetching the page, running the deterministic checks, keeping a whole page in one pass, and refusing to show you a fix that points at an element that is not in your markup.

The four steps

What the model is actually asked

01routed to gpt-4.1-mini

Read the page markup

The cleaned markup, scripts and styles stripped, in one pass where it fits, with a system prompt that names the barrier families that actually block people: labels and names, keyboard operability, structure and headings, contrast and meaning carried by color, alt text, and wrong ARIA. The model is told to quote the offending element verbatim and to leave out anything it cannot quote exactly.

02routed to gpt-4.1-mini

Rank each barrier by impact

Every issue is rated critical, serious or minor by how much it blocks real people on this page, and the list is ordered so the thing to fix first is at the top. An issue that is genuinely minor is rated minor, because a report that flags everything as critical is a report nobody reads twice.

03routed to gpt-4.1-mini

Explain who it affects, plainly

Two to four sentences naming the people and the consequence on this page: the screen reader user who hears only the file name, the keyboard user who cannot reach the control at all. No jargon the reader has to look up. The test we hold it to: could a developer act on this without reading the spec.

04routed to gpt-4.1-mini

Draft the code fix

The corrected markup for the reader's own content, minimal and faithful, with a one-line note on why it works. This is a separate routed step because auditing and fixing are different jobs and may not want the same model for long.

Routing

The table, generated from the code

This is not a diagram somebody drew. It is rendered from the same routing table the audit engine reads, so if it is wrong here it is wrong in production.

StepModelProviderTierUSD per M tokens
Read the page and find the real accessibility barriersgpt-4.1-miniopenaibalanced$1.60
Draft the exact code fixgpt-4.1-miniopenaibalanced$1.60
Write the overall read for the reportgpt-4.1-miniopenaibalanced$1.60
Work out what kind of page this isgpt-4.1-nanoopenaifast$0.40

Escalation

Above roughly 30,000 characters of markup the audit escalates to the frontier model rather than the balanced one, because that is where whole-page accuracy starts to matter more than cost. Audits are capped at 48,000 characters in one pass, and anything past that is reported rather than dropped silently.

Fallback

If the routed model fails or returns nothing, the call walks a chain of candidates from other tiers and other labs before giving up. One provider having a bad afternoon should not cost a customer their audit, and it does not.

Verification

Every quoted snippet is located in your own markup before it is shown, and the structure is validated against a schema. A snippet that does not appear is dropped and counted. This is the one part of the pipeline the model does not get a vote on.

Candidates

What is wired, and what is one key away

Every model below sits behind the same interface. Adding a lab is one case in one file, which is the entire point of building it this way.

gpt-4.1

frontier

openai · 1,000,000 token context

  • large, script-heavy pages
  • issues that depend on the whole page structure
  • unusual ARIA and custom widget patterns

gpt-4.1-mini

balanced

openai · 1,000,000 token context

  • the main page audit
  • plain-language explanations and impact
  • drafting code fixes

gpt-4.1-nano

fast

openai · 1,000,000 token context

  • page-type triage
  • deduplicating near-identical issues

claude-sonnet

frontier

anthropic · 200,000 token context

  • careful reading of complex markup
  • fix wording and tone

gemini-flash

fast

google · 1,000,000 token context

  • bulk triage across a whole-site crawl
  • cheap re-checks

llama-open

open

meta · 128,000 token context

  • self-hosted audits for teams that cannot send markup out

In this deployment the OpenAI and Anthropic adapters are written and the OpenAI one is live. The rest are declared with their real model names and switch on when their key is present. No model is trained on the pages you scan, by us or by our providers.