Semantic Compute: From Interpreters to Compilers
Jev's rapid adoption suggests that developers are ready for a different way to execute intelligence inside software. We believe it points toward a larger transition: from calling models to compiling semantic functions.
Jev, from TypeSafe AI, takes context and questions and returns typed decisions: choices, scores, and yes-or-no assessments. Its reception has been striking. On September 18, Vercel reported that Jev had reached nearly 13% of AI Gateway's paid teams within 24 hours, more than twice the first-day adoption of any previous model launch on the platform. That is a platform-specific launch signal, with retention still to be established, but it is strong evidence of developer interest in an interface organized around decisions. Vercel's launch report makes that distinction worth taking seriously.
The integration pattern is equally revealing. Browserbase describes adapting Stagehand so Jev classifies actions and selects among candidates assembled by ordinary code, with an LLM fallback when the selection is rejected. Its early tests reported median Act latency falling from 1.97 seconds to 0.46 seconds. The implementation was manually engineered, and those measurements belong to that experiment. Nevertheless, it demonstrates the architectural move: expose the structure of a task, assign its parts to suitable execution mechanisms, and preserve a more general path for the remaining cases. Browserbase's implementation gives the broader thesis a concrete foundation.
Our inference is that these are early signs of Semantic Compute becoming a useful programming paradigm: represent semantic behavior independently of the model that currently performs it, then optimize how that behavior executes. Frontier LLMs provide extraordinarily flexible first implementations. Jev, GLiNER, retrieval systems, conventional ML, databases, and deterministic code enlarge the space of possible replacements. A compiler and runtime could make that space accessible through a stable application abstraction. This is the upcoming direction we are exploring for Seldon. The architecture below is a proposal; Jev's traction supports its premise, while automatic compilation, trustworthy evaluation, and useful economics remain hypotheses to test.
Frontier models are semantic interpreters
A substantial class of LLM calls begins and ends in ordinary program state. A support ticket becomes a queue; an invoice becomes a vendor ID; a document becomes a structured record; a collection of passages becomes a ranked list. The application supplies a task description, data, and constraints. A general-purpose model interprets them and produces tokens that the application converts back into values. This is an unusually flexible execution interface: developers can express unfamiliar concepts, policies, exceptions, and relationships before they have designed features, collected a training set, or implemented a dedicated algorithm.
Calling the model a semantic interpreter describes that role in the software stack. In the expression LLM(P, x), P is a relatively stable task description and x is changing application state. Once a workload repeats, we can ask whether a specialized f_P(x) would implement the intended behavior more efficiently. Partial evaluation supplies a classical precedent for specializing a program with respect to known inputs. The analogy has a precise limit: replacing an LLM with a classifier usually approximates a task rather than derives a behaviorally equivalent residual program. Even perfect reproduction of the original model would preserve its errors. The specification we ultimately care about belongs to the application.
The opportunity follows from separating logical behavior from physical execution. An operation that resolves an entity may execute as a lookup, a retrieval-and-reranking pipeline, a specialized decision model, or a frontier-model call. Those implementations differ in applicability, cost, latency, and error distribution. A bounded result does not imply an easy problem—a boolean decision may require substantial reasoning—but it reveals structure that a general text interface can obscure. Semantic Compute makes that structure explicit enough for software to inspect and optimize. Its novelty would lie in making heterogeneous semantic execution a managed systems layer, building on decades of classifiers, databases, compilers, and ML pipelines.
The durable object is a semantic function
The application should depend on a Semantic Function: a typed mapping together with the conditions under which its behavior is acceptable. Consider a vendor resolver:
def resolve_vendor(invoice: Invoice, catalog: VendorCatalog) -> Resolution:
return current_llm_resolver(invoice, catalog)
# Resolution = Matched(VendorID) | Unresolved(reason)
The signature is only the structural part of its contract. Behaviorally, it must identify the issuing legal entity, return an ID from the supplied catalog, and allow an unresolved result when the evidence cannot distinguish candidates. Operationally, the application may impose a latency budget, limits on false matches, and different requirements for particular suppliers. These requirements have distinct authorities: types describe permissible values; business policy defines acceptable behavior; evaluation estimates performance on a stated population; runtime budgets constrain execution. Combining them into one confidence threshold would erase information the optimizer needs.
This boundary can survive a complete change of implementation. A resolver might begin as a prompted frontier model, become span extraction followed by candidate retrieval and selection, and later use authoritative identifiers for a large subset of inputs. The caller still consumes Resolution. The interesting abstraction is therefore the function and its contract, while the execution plan becomes replaceable. A developer should be able to supply an existing function, representative examples, and consequential policy decisions without first learning a new instruction set. The compiler's vocabulary can be much more explicit than the product's adoption interface.
The distinction also clarifies the authority of production traces. Traces reveal which inputs occur, which implementations run, what they cost, and where disagreements concentrate. They are profiles and observations of existing behavior. They cannot decide whether a parent company may substitute for a subsidiary, whether a missing vendor should be created, or which error is unacceptable. Those decisions must come from application policy, trusted outcomes, or developer judgment. Inference can propose a contract; responsibility for its meaning remains visible.
A kernel for machine-facing semantics
A useful intermediate representation should preserve the kinds of values software actually consumes. We propose six candidate operation families: PREDICATE, CHOICE, POINTER, SPAN, SCORE, and GENERATE. Their purpose is practical: expose meaningful constraints, support multiple implementations, and retain information useful to optimization. This is a candidate kernel, not a proof of mathematical minimality or universal expressiveness. A small list can be expressive by hiding arbitrary computation in an operation; that says little about whether its programs are analyzable, optimizable, or affordable to evaluate.
PREDICATE assesses a proposition about an input, such as whether a passage supports a claim. CHOICE selects from a declared finite set, such as a support queue or a next action. A predicate can be encoded as a two-way choice, but retaining proposition semantics is useful for decision thresholds and asymmetric error costs. Choice preserves a closed output domain: an implementation must return an allowed alternative or abstain. This property can be enforced for generative and discriminative backends alike; it constrains the answer's shape without establishing that the answer is right. Jev's Choice interface is one physical realization of this logical operation.
POINTER selects the identity of an existing object, while SPAN identifies locations in a source. A pointer should carry enough context to establish which collection and version its ID belongs to; a span should identify the source version and offset convention, with multiplicity and overlap defined by the task. These are richer contracts than strings. Returning a catalog reference avoids reconstructing identity from a generated name. Returning a source range preserves provenance through later parsing and normalization. Although a span could be encoded as a pointer over substrings, retaining its geometry and source relationship enables more useful implementations and checks.
The invoice example exposes the boundary. “Acme” is a valid source span, and both Acme Industrial Services and Acme Software may be valid catalog references. Neither fact settles which legal entity issued the invoice. Source grounding establishes where evidence came from; candidate membership establishes where an answer may point. Semantic correctness still depends on the relationship between that evidence and the task. A retrieval stage can also exclude the correct entity before selection begins. GLiNER, which identifies spans under supplied labels, and GLiNER2, which extends the interface to classification and structured extraction, make useful targets for parts of this program.
SCORE evaluates an input under an explicit rubric and scale. An ordinal urgency assessment and a continuous relevance estimate have different semantics; neither automatically supplies a probability of correctness. Jev's Score primitive, specifically, evaluates ordered descriptive levels. GENERATE produces an open-ended sequence when the required result is language, code, or another sequence whose content cannot usefully be selected from a supplied domain. Generation remains a first-class operation with its own budgets and validation requirements. The architectural change is to make it an explicit choice of computation, rather than the implicit representation of every semantic task.
The kernel does not replace ordinary programming. Records, collections, arithmetic, iteration, branching, and effects still need host-language or execution-IR semantics. Relational operators such as selection, projection, join, and aggregation describe how collections are transformed; semantic operations describe judgments used within those transformations. A join predicate might be exact equality or a learned relation, with very different guarantees. General computation can come from the surrounding language, but Turing completeness cannot establish that a model implements an intended judgment or that a rewrite preserves it. Those are specification and validation problems.
Preserve abstraction until lowering becomes useful
A practical compiler should retain four conceptual levels: domain functions, a high-level semantic IR, the kernel, and physical execution plans. resolve_vendor retains business meaning. EXTRACT and MATCH expose the structure of the task. SPAN and POINTER make provenance and identity explicit. The physical plan binds that structure to parsers, indexes, model versions, and an acceptance policy. This hierarchy is a repertoire of representations, rather than a mandatory sequence through which every operation must pass. A backend capable of joint structured extraction may implement a high-level operation directly.
Early decomposition can destroy valuable information. EXTRACT should retain a record schema so an optimizer can consider joint extraction and cross-field constraints. RANK should retain list-level semantics because pairwise comparisons or a dedicated reranker may outperform independent scores followed by sorting. MATCH preserves opportunities for normalization, candidate blocking, and indexed lookup. FILTER exposes collection-level batching; COMPARE preserves a relation between inputs. These operators may lower into the kernel, remain intact for a specialized backend, or fuse with adjacent operations. The criterion is whether lowering exposes a better executable plan while preserving the relevant contract.
MLIR supplies a useful precedent for multiple abstraction levels and progressive lowering. Halide demonstrates the value of separating an algorithm from its execution schedule. There is also direct prior art in AI systems: Palimpzest explores declarative AI workloads and physical plans with cost, runtime, and quality tradeoffs; DSPy optimizes declarative model programs. These systems make the compiler direction credible and constrain any novelty claim. The proposed Seldon problem is to make heterogeneous optimization usable around existing application functions, including their qualification and maintenance as workloads and backends change.
Jev and GLiNER occupy the physical side of this separation even though their APIs also expose useful logical operations. Both still consume task information at runtime. Specialization need not eliminate interpretation completely to be valuable: it can change the architecture, output mechanism, or amount of computation used for a recurring judgment. Keeping the logical IR independent of either model makes room for a customer-specific classifier, a new extractor, or an exact program to win later. This is what allows the optimizer to cross computational paradigms as well as model providers.
Optimize programs, including which calls disappear
Backend selection is only one dimension of the search. A compiler can change decomposition, candidate construction, batching, caching, and control flow. In the vendor example, a trusted identifier may make semantic matching unnecessary for a subset of invoices. If the contract recognizes that identifier, its provenance is valid, and it uniquely resolves within the applicable catalog, the compiler can propose a guarded lookup. The same rewrite would be unjustified for arbitrary text that merely resembles an identifier. An alternative plan might fuse extraction and resolution into one learned operation, trading intermediate observability for a potentially cheaper execution path.
These transformations require different kinds of justification. A behavior-preserving rewrite, such as reusing a pure deterministic computation with identical inputs and dependencies, follows from semantics. Conditional specialization, such as the authoritative-identifier lookup, depends on a contract and checked preconditions. A learned substitution, such as replacing an LLM with an extractor and classifier, requires empirical evidence about changed behavior. The compiler should attach those obligations to its candidates rather than treating every change as an interchangeable optimization. Whole-program search matters because independently choosing the cheapest model at each node can miss a plan that removes those nodes altogether.
The full optimization target is an execution policy, including the conditions under which each implementation runs:
Given candidate policies Π and workload population D:
choose π ∈ Π to minimize E[x ~ D][runtime_cost(π, x)]
subject to declared limits on:
critical-error probability, including required workload slices
unresolved-result probability
end-to-end p95 latency
π includes the whole conditional program:
retrieval + backends + applicability checks + acceptance + fallback
Each quantity needs an operational definition. A false match among all requests differs from a false match among accepted matches; both may matter. A risk limit must identify its population and evaluation period. Cost includes rejected specialist attempts followed by fallback, and tail latency includes requests that traverse multiple paths. The search also has a budget: generating candidates, collecting labels, evaluating them, and maintaining a selected plan consume resources. The economically best result can be a single general model, a small conditional program, or retaining the current implementation.
Uncertainty belongs in the contract and runtime
A typed semantic value needs an explicit path for insufficient evidence. A conceptual result type might be:
SemanticResult<T> =
Answer(value: T, provenance, assessment?)
| Abstain(reason)
assessment = backend-specific scores or distributions,
interpreted under a versioned acceptance policy
This representation separates an answer, its evidence, and a model's assessment from the decision to accept it. A deterministic lookup need not invent a confidence score. A model distribution may be useful without being calibrated for the application's traffic. A rubric score has different meaning from a probability, and neither should silently become a universal measure of trust. For example, Jev's Choice confidence is derived from the top option probability and the number of alternatives. It is not an observed accuracy rate on the customer's vendor-resolution task. Acceptance thresholds require evidence for the actual candidate construction, workload, and error categories.
A tracing JIT offers a productive analogy: observe common execution, specialize under assumptions, and return to a general implementation when those assumptions fail. Semantic specialization adds a fundamental complication. An unsupported document format or stale catalog version can be detected exactly; a plausible but incorrect entity selection may pass every applicability, schema, and confidence check. Application traces profile externally visible behavior, rather than exposing the model's internal execution trace. The frontier model can provide a useful general fallback, but its presence does not turn statistical acceptance into semantics-preserving deoptimization.
A fallback cannot fix an error the system accepts. Selective prediction provides relevant machinery for studying the relationship between acceptance coverage and errors among accepted predictions. A production policy must additionally account for the fallback's errors and unresolved results. Rejecting almost everything can improve a specialist's accepted accuracy while creating little economic value. Conversely, increasing acceptance may save inference spend while exposing precisely the ambiguous inputs on which an application's losses concentrate. The object of evaluation must therefore be the complete policy, with enough detail to distinguish these outcomes.
Effects require an equally explicit boundary. The semantic kernel should compute judgments and values; ordinary application code should own authorization, transactions, and externally visible actions. Selecting Refund does not authorize a payment. Selecting a tool does not execute it. This separation makes candidate replay possible without repeating customer actions and allows existing workflow systems to retain control flow. Reproducible evaluation also requires captured or versioned dependencies: a resolver reading today's changing catalog cannot be assumed to replay yesterday's computation. Caching stochastic outputs likewise changes their joint behavior and needs an explicit application policy, rather than an assumption that all model calls are pure deterministic functions.
Statistical compilation needs an evidence discipline
The proposed compiler would combine exact transformations with empirical qualification. For the latter, a finite dataset estimates behavior under assumptions about how examples represent future requests. It cannot certify arbitrary future inputs. Development data used to construct candidates and tune thresholds must be separated from final evaluation; repeated optimization against a holdout makes that holdout part of development. The evaluator itself needs an authority model. Historical outputs can contain mistakes, and a frontier-model judge can share a candidate's errors. Developer-confirmed examples, invariants, and observed outcomes supply different forms of evidence, each with limitations that should remain attached to a qualification result.
For vendor resolution, evaluation should distinguish correct matches, false matches, and unresolved cases across ambiguous subsidiaries, missing vendors, unfamiliar aliases, and catalog changes. Near-duplicate invoices must not leak between development and final evaluation. Retrieval and selection must be evaluated together: no downstream classifier can choose an entity excluded from its candidate set. Rare critical failures may require more representative evidence than a team can afford to collect. In that case, the correct output of compilation may be a narrower applicability region or a report that the proposed risk target is unsupported. A precise numerical target becomes useful only when paired with a procedure capable of assessing it.
Qualification also expires under relevant change. Model versions, prompts, schemas, retrievers, catalogs, and acceptance policies belong to the executable plan's identity. A changed dependency or workload can invalidate earlier evidence, even when the function's type remains stable. The runtime therefore needs observable execution paths, bounded rollout, rollback, and a way to revisit assumptions. Production observations can trigger investigation and new evaluation; they cannot automatically establish correctness when the true outcome is unobserved. This lifecycle is where a compilation system could create sustained value beyond a one-time model substitution.
A new execution layer for AI software
The larger bet is that machine-facing semantic computation has enough recurring structure to support a durable logical layer. Today, choosing a model often implicitly chooses an algorithm, an output representation, an uncertainty interface, and a cost structure at once. Semantic Compute would let an application name the behavior it needs while infrastructure searches those choices separately. A function could retain its identity across a frontier model, a specialized model, a retrieval pipeline, and ordinary code. That separation would make improvements in any execution family available to software whose authors never anticipated the particular backend.
Seldon's proposed role is to make this transition practical: begin with an existing function, establish its behavioral contract, profile its workload, construct alternative programs, and show a developer the tradeoffs and consequential disagreements before changing execution. The kernel is the compiler's vocabulary; the function is the developer's boundary. A useful product must reduce the work of establishing and maintaining a trustworthy replacement, including the cases where the best answer is to keep the existing implementation. Time to a justified replacement matters alongside acceptance coverage, complete-system errors, unresolved outcomes, cost, and latency.
This thesis has clear ways to fail. General models could become inexpensive enough that specialization rarely repays its overhead. Evaluation could cost more than the saved execution. Manual replacements could be sufficiently easy and stable that another systems layer adds little value. Existing frameworks or providers could absorb the necessary optimization. Jev's launch does not settle those questions. What it does supply is timely evidence that developers will adopt a specialized semantic interface, while integrations such as Stagehand show that useful programs can span models and ordinary code. The next experiment is whether a compiler can make those programs easier to produce and trust.
We believe the resulting paradigm is worth building toward: semantic intent becomes a stable software object; its implementation becomes an evolving execution plan. Frontier LLMs make the first version possible. Specialized models expand the available targets. A compiler and runtime would connect the two, turning each new execution mechanism into another way to implement existing application behavior.
Have you replaced a production LLM call with rules, a classifier, an extractor, or a smaller model? Share the function, one difficult input, and what took the most work to validate. If you reverted, what failed? For compiler and ML engineers: which representation, transformation, or evidence requirement has the weakest foundation? Reply where you found this essay or send us a private note; anonymized examples are welcome.
Help shape what comes next
This thesis describes an upcoming direction for Seldon. Share an LLM-powered function you would like to replace, a failed attempt, or a technical objection. Anonymized examples are welcome.