How it works

Contracts define behavior. Traces provide evidence.

Today, Seldon routes LLM calls, analyzes recurring work, and provides tools to inspect built compiled paths. Its hybrid compiler uses a bounded, schema-typed program format. Our upcoming direction extends this foundation toward optimizing existing application functions across ordinary code, specialist models, and LLMs.

Read the thesis and help shape the upcoming direction →
01

The thesis

Some LLM calls contain familiar data operations.

Classification, extraction, matching, and arithmetic can appear together inside one LLM call. Making those operations explicit can reveal alternative implementations using ordinary code and learned models. Whether an alternative is useful depends on the behavior the application needs and the workload it receives.

In the upcoming direction, the starting point would be an existing function and its behavior contract. Routed traffic is one source of examples and performance profiles. It can suggest opportunities, but does not fully specify what the function should do.

02

Behavior contracts

Make the intended behavior explicit.

Code, types, prompts, tests, and developer-confirmed policies help define a function. Traces contribute examples and reveal recurring work. These sources need to be considered together, especially when a historical answer conflicts with a business rule.

Intended behavior

The function the application needs, its inputs and outputs, and the business rules that distinguish a correct result from a plausible one.

Domain context

The entities and relationships that matter to the decision. Traces can suggest this context, but cannot establish whether two vendors, accounts, or subsidiaries are interchangeable.

Evaluation evidence

Representative examples, expected outcomes, unresolved cases, and important errors. Historical model answers are a baseline to compare with, not an independent source of truth.

Seldon already uses a bounded, schema-typed execution graph: operations have explicit inputs, parameters, and output schemas. The proposed direction would preserve higher-level semantic operations before lowering them into a typed intermediate representation and choosing their implementations. Types help validate composition; they do not prove a learned judgment is correct.

Fig.1 — Behavior and evidence to a candidate program

01behavior contract+ traffic evidence02opsidentify operationstask vocabulary03typed operatorscandidate program04evaluatecandidatefrontier LLM · fallbackdetected limitations
03

Program synthesis

Compose explicit operations, then evaluate the complete program.

The current hybrid runtime executes bounded programs containing deterministic operations and an optional model call. The thesis explores how a compiler could choose among a broader set of implementations for an existing function, including specialist models. This broader function-optimization experience is an upcoming direction.

The task families below are useful vocabulary for describing work and finding candidate implementations. They are not the execution language itself. The typed program makes data dependencies and operations explicit, while evaluation tests whether a candidate satisfies the intended behavior on representative inputs.

A

Text classification

text → label ∈ a small fixed set

B

Sequence / span labeling

output spans ⊂ the input text

C

Structured extraction

output is a stable JSON schema

D

Entity matching / retrieval

output references a bounded catalog

E

Similarity / pairing

pairwise comparison or grouping

F

Normalization / transform

deterministic function of input

G

Numeric / analytical

a number derivable by computation

An illustrative candidate might extract supplier evidence, look up vendor records, and select a match or return unresolved. Another might keep the original LLM call. Model calls can remain at runtime, and changes in inputs or catalogs can require reevaluation. The useful comparison includes quality, accepted coverage, cost, and fallback latency.

How to assess a replacement

Evaluate the behavior. Understand the limits.

01

Agreement is one measurement.

Recorded-answer replay measures whether a candidate reproduces earlier outputs. Correctness needs evidence tied to the intended behavior, including cases where the original model was wrong.

02

Functional checks and quality answer different questions.

A program can execute successfully and return the right schema while making the wrong decision. Evaluation should examine errors, accepted coverage, cost, and latency for the complete program.

03

Specialization can cover part of a workload.

A specialized path may handle only supported inputs. Model confidence needs workload-specific evaluation; component confidence values do not automatically establish end-to-end quality.

04

Fallback handles detected limitations.

Unsupported requests and execution failures can return to the original provider path. A confidently wrong result can still pass the checks, so fallback is not a correctness guarantee.

See it on your own traffic.

Start with routing and Live Audit to inspect recurring work and possible replacements. Read the thesis to explore the upcoming function-optimization direction.

How it works | Seldon