Skip to content
Nate G. / Vancouver, WA · US
all work
Open SourceTypeSafe AIJevSystem OneSemantic Evidence

Agent-Workflow TypeSafe AI

Bounded TypeSafe AI/Jev semantic evidence for Agent-Workflow without transferring lifecycle authority.

Case summary

Purpose

Some engineering decisions are semantic rather than purely rule-based, but that does not require handing workflow authority to a model. This adapter turns a narrow set of judgments into typed, auditable evidence while the host keeps policy, execution, review, and acceptance deterministic.

System

Agent-Workflow TypeSafe AI is the standalone TypeSafe AI/Jev adapter originally built for Agent-Workflow's plugin boundary. It projects bounded task or skill evidence into Choice, Noul, and Score questions, normalizes the results into versioned semantic receipts, and returns advisory evidence to host-owned policy.

Contribution

Semantic-vs-deterministic authority boundary, projection/receipt/fallback/compatibility model, and expansion into Agent-Workflow's built-in provider, comparative evaluation, and benchmark evidence.

Architecture & engineering model

Bounded semantic evidence pipeline

flow

Semantic evidence is projected and normalized before host policy can use it; model output never directly becomes a workflow action.

application evidence
select / bound / redact
Choice / Noul / Score
Jev / System One
normalized semantic receipt
application-owned policy
action / advisory / escalation

Integration evolution

evolution

The provider seam was first proven as a standalone adapter, then moved into Agent-Workflow for the current runtime while comparison semantics were extracted into a provider-neutral library.

Standalone adapter

  • provider transport and compatibility boundary
  • bounded projection and typed question sets
  • semantic receipts and explicit fallbacks

Built-in Agent-Workflow provider

  • current optional routing provider
  • host-owned policy and lifecycle
  • used by current benchmark qualification paths

Provider-neutral evaluation

  • agent-workflow-comparative-eval
  • versioned observations and frozen corpora
  • pairing, cohort identity, metrics, deterministic statistics

Engineering principles

Confidence is evidence, not authority

Probabilities, distributions, and confidence can inform policy; they do not grant permission to mutate state or bypass deterministic controls.

Shadow before promotion

Candidate semantic decisions can be captured comparatively before host policy is changed to consume them.

Failure stays explicit

Uncertainty, unavailable credentials, SDK failures, and policy rejection resolve to named statuses and deterministic fallback paths.

Request fidelity is part of semantic assurance

A valid provider call is not enough: authoritative evidence projection is versioned, hashed, and constructed deterministically when the task requires exact source fidelity.

Problem

  • Pure rules are reliable for explicit invariants but can be brittle at bounded semantic seams.
  • Semantic evidence needs provenance, redaction, uncertainty handling, and deterministic fallback before it can safely influence application behavior.
  • A provider integration should not become workflow, routing, review, or acceptance authority.

Technical Highlights

  • Bounded and redacted request projection
  • Typed Choice, Noul, and Score question sets
  • Versioned semantic receipts with request hashes and source references
  • Explicit uncertainty, unavailable-service, and no-action fallbacks
  • Credential-free compatibility diagnostics
  • Optional comparative-evaluation integration
  • Agentic-Jev request-fidelity work now separates generic request mechanics from task-specific completeness policy; the v2 SWE-manager path injects exact authoritative task/proposal context in deterministic host code

System Design

  • task or skill evidence → validate / project / bound / redact → typed questions → optional TypeSafe System One call → normalized receipt → Agent-Workflow policy
  • The standalone adapter owns provider transport and evidence normalization; Agent-Workflow owns policy and lifecycle authority.
  • Agent-Workflow 0.11.x also contains a built-in optional routing provider; current benchmark studies use that built-in provider rather than this compatibility adapter.

Deterministic Authority

  • Question-set and projection validation
  • Secret redaction and payload bounds
  • Receipt schema validation
  • Host-owned thresholds, fallback, routing, execution, review, and acceptance

Where Models Fit

  • Resolve explicitly registered bounded semantic questions through TypeSafe AI/Jev.
  • Return typed evidence; never directly perform host workflow actions.

Trust, Failure & Recovery

  • Missing credentials, SDK availability problems, uncertainty, and provider failures resolve to explicit statuses rather than implicit authority.
  • The package publishes an exact compatibility boundary and does not infer compatibility from repository activity.
  • Optional API-call logging is off by default and treats retained projected text as sensitive.

Measured evidence

Related built-in provider qualification

measured

9/9 benchmark phase qualification calls completed across BM3–BM5

The public benchmark program records three pre-treatment TypeSafe routing qualification calls per study across BM3, BM4, and BM5.

Current benchmark studies use Agent-Workflow's built-in provider, not this standalone compatibility adapter. The qualification evidence demonstrates the bounded question model in the current host integration; it is not a direct performance validation of plugin 0.1.2 and does not establish a causal TypeSafe effect.

n=92026-09-23source fafe862c0aa3
evidence method/source

Evidence & documentation