Completion needs an acceptance boundary
A worker finishing does not establish that the requested work is complete. Execution, evidence, review, and acceptance need distinct states, with explicit rules for accepting the result.
Completion gatesProfessional profile
Agentic Systems / AI Platform Engineer
I build reliable autonomous execution systems with explicit boundaries around authority, evidence, evaluation, recovery, and accountability.
Built on more than a decade of production backend and systems engineering.
Vancouver, Washington · US-based · Open to senior US-based Agentic Systems / AI Platform engineering opportunities.
Software / systems / agents
Recent work on agent-assisted software engineering: making context, specifications, execution, judgment, and evaluation explicit enough to inspect and improve.
The Work Atlas records the systems, studies, failures, corrections, and evidence behind that work.
More than a decade of production software engineering across backend services, APIs, databases, distributed systems, cloud and CI/CD, greenfield development, brownfield systems, and legacy modernization. That background is the baseline for the recent agentic work below.
.NET / C#production backendIndependent tools connected through specifications, contracts, orchestration, evidence, and evaluation.
Source, structure, and codebase context.
Requirements, decisions, and durable baselines.
Isolation, recovery, review, and acceptance.
Bounded judgment with explicit uncertainty.
Comparison, produced software, and inspectable evidence.
Capabilities are facets. A project can appear in more than one; a shared concern does not imply a dependency.
Numbered arrows show documented relationships between independent projects; they do not imply common ownership. Dashed lines indicate optional integration. Select a node to inspect its connections.
No project selected
Each project began with a concrete engineering problem. The project, decision, and evidence links show how the answer changed as the work matured.
Current principles and evidence records stay inspectable without turning the page into a vertical document wall.
A worker finishing does not establish that the requested work is complete. Execution, evidence, review, and acceptance need distinct states, with explicit rules for accepting the result.
Completion gatesRecovery and review mechanisms can add overhead to routine work. Preserve the boundaries that matter, simplify the common path, and measure the consequences.
Simplification decisionsP3 review caught that the original publisher would have copied private oracle provenance. The published result therefore uses a separately defined public projection with explicit allowlist, integrity, sample-identity, and privacy verification.
Publication integrityA Codex-LB shutdown investigation deliberately separated code reading, empirical reproduction, an independent verifier using a different method, and orchestrator remeasurement. Later upstream review found an additional bounded-wait risk and replaced a merely correlated regression trigger with a deterministic barrier. The upstream PR remains under review.
Upstream Codex-LB reviewRecent Agent-Workflow use exposed four narrow contract failures: stale closeout guidance, equivalent cwd evidence scored as contradictory, recursive scope expressed as an exact path, and diagnostic loss when the configured state root was unavailable. Each became a bounded merged correction without rewriting the historical failure into a broader claim.
Diagnostic preservation correctionMore context can make a semantic answer more decisive without making it more discriminating. Primary artifacts and deterministic tool output should be separated from agent-authored claims; prior semantic answers and the choosing agent’s verdict should stay out of the independent provider projection by default.
Neutral projection rulesA bounded, typed, successfully dispatched semantic request can still omit the authoritative evidence needed for the decision. In the first fresh manager study, 0/20 successful Jev calls met the frozen completeness check even though the bridge preserved every dispatched string. Request construction therefore became a deterministic, versioned part of the treatment.
SWE-manager v2 request policyTyped semantic judgment returns bounded evidence while application code retains policy, side effects, lifecycle, review, and acceptance authority.
The integration explores bounded semantic judgment through typed Choice, Noul, and Score evidence. Receipts preserve provenance: source references, question versions, model details, request hashes, and outcome status. Uncertainty, no-decision outcomes, and fallback remain explicit.
Application code determines how evidence affects the result. TypeSafe/Jev does not own orchestration, side effects, review, or acceptance, and can be used independently of Agent-Workflow.
Structural checkpoint update. Controlled skill qualification did not translate into reliable optional uptake: six frozen real manager decisions made zero Jev calls. Agent-Workflow therefore prototypes a provider-neutral outer checkpoint that records the choosing agent’s draft before semantic evidence, projects neutral primary evidence without that verdict, and requires explicit reconciliation on disagreement or uncertainty.
Context discipline. Later review work reinforced that richer context can make semantic output more decisive without making it more discriminating. The independent projection now separates deterministic tool evidence from agent-authored claims and excludes the choosing agent’s preferred answer and prior Jev answers by default.
Fresh manager effectiveness study. A 30-pair SWE-Lancer manager study then observed 15/30 correct in both control and treatment (0.0 pp difference; paired 95% bootstrap interval -16.7 to +20.0 pp). Jev was invoked on 20 treatment attempts, but 0/20 successful calls met the frozen context-completeness check. The result is preserved as evidence for the treatment that actually ran rather than retroactively repaired.
Request-fidelity correction. The audit showed that the bridge did not truncate the successful calls; the coding agent had compressed authoritative task and proposal material before dispatch. The next treatment generation therefore builds the neutral request in deterministic host code. That v2 path has passed live qualification on an already-observed task; its fresh effectiveness cohort is preregistered but unrun.
Evidence boundary. Provider success alone does not establish correctness. The completed routing-semantic-v1 study now supplies independent seam-level correctness evidence for a frozen 120-case cohort; those findings should not be generalized into universal Jev superiority or downstream software-quality causality.
One retained development pair used approximately 185 seconds less wall time and 235,308 fewer provider tokens with the slimmed Agent-Workflow treatment, while scoring 2.5 machine-score points lower.
Boundary: one development pair, no human-complete pair, and no measured cost saving.
Paired benchmark runs retain the produced software alongside timing, token usage, scoring, and methodology. This makes changes to the process inspectable against an explicit control.
This was one development pair and is not a generalized performance claim.
The pair is awaiting human review; no human-complete pair is recorded. Cost evidence is unavailable, so fewer tokens should not be presented as a measured cost saving.
TypeSafe qualification ran before both treatments, with semantic routing outside the comparison. These results do not establish a TypeSafe effect.
The study moved from asking whether semantic routing behaves differently to asking whether it is actually more correct against independently frozen ground truth.
Current boundary: v1 is published. V2 achieved a full 0.6.3 qualification pass, but the 0.6.4 model-facing schema representation changed and must qualify again before real A/B. No real v2 A/B/C result is claimed.
The methodological shift is the important part of this work. Early instrumentation could show that TypeSafe/Jev and deterministic routing sometimes behaved differently. That was not enough to say which decision was better.
Independent ground truth is now frozen. The qualified DeepSeek-Flash adjudicator cohort completed independent A and B labeling on the same hash-pinned blinded view, C received only disputed cases and seams, and genuine three-way conflicts went through recorded human resolution. The resulting oracle was frozen and validated before any live TypeSafe/Jev comparative inference.
Oracle construction now has a recorded cost/usage baseline. The initial qualification and adjudication runs through P0B oracle freeze used the OpenAI Codex harness, routed through codex-lb to the DeepSeek-4.1-Flash API. Provider accounting for that baseline was $0.07 USD across 36 API requests and 726,263 tokens. These totals stop at oracle freeze and do not include P1/P2 Jev traffic.
The staged execution path is now complete. P1 first verified the live evidence path on 8 development-only cases. P2 then ran all 120 frozen cases without exposing the oracle during inference, producing 360 observations from 120 provider requests with 0 exclusions. P3 independently verified the sanitized publication projection before the result was published.
The result is deliberately heterogeneous. In this frozen cohort, the semantic candidate was more accurate on task-class and interaction-required classification, while semantic-risk ordinal error was approximately unchanged at the study's resolution. All 120 provider requests succeeded. Provider treatment-cost evidence is incomplete.
Raw semantic quality and operational policy remain separate layers. Choice, Noul, and Score evidence are evaluated against the oracle without treating the current confidence threshold, fallback rule, or non-automatable risk policy as model correctness.
Decision-provenance correction. The authoritative v1 A/B/C passes retained labels but no structured per-decision justifications, so the original human resolver worked from the case, metadata, rubric, and labels. A later audit found contemporaneous provider reasoning summaries in private Inspect logs. Those summaries were supplementary execution evidence, were not available to the resolver, are not hidden chain-of-thought, and are not being retrofitted into the frozen oracle. The failure was promotion of useful evidence into the durable decision record, not complete historical loss.
Reporting correction. P2 exposed one 0.99-total probability vector in each multiclass seam. Raw evidence was preserved; only the derived calibration copy was normalized, and the deviation was disclosed.
Publication correction. P3 review found that the first publisher would have copied the private oracle verbatim. The released bundle instead uses a redacted public oracle projection and passed an independent allowlist, manifest-integrity, cohort-identity, and privacy verifier before publication.
Results boundary. Routing Semantic v1 is complete and P3-published: 120 frozen cases produced 360 decision observations from 120 semantic provider requests with 0 exclusions. In this cohort, semantic-candidate accuracy was 81.67% vs 52.50% for task class and 91.67% vs 71.67% for interaction required; both paired 95% accuracy-difference intervals remained above zero. Semantic-risk MAE was 0.28975 vs 0.29167 and its paired interval spanned zero. All 120 provider requests succeeded. Provider treatment-cost evidence is incomplete, so no cost-advantage, universal-superiority, or downstream software-quality claim is supported. Routing-semantic-v2 later reached a first full live IA-1 through IA-11 qualification pass under Benchmark 0.6.3 with sanitized proof that strict JSON-schema controls reached Codex-LB ingress. Before any real v2 label was produced, real-cohort readiness review found that the 120-case model-facing schema expanded to 225,971 serialized characters because an identical record shape was repeated per case. Benchmark 0.6.4 replaces that representation with the eligibility-grouped-case-enum/v1 strategy, reducing the equivalent schema to 3,310 characters while preserving all 120 case identities. Because the model-facing schema changed, the earlier pass cannot authorize the 0.6.4 cohort: a fresh live IA-1 through IA-11 qualification is required. Real routing-semantic-v2 A/B/C remains unrun. Separately, a fresh 30-pair Agentic-Jev SWE-manager v1 study completed with 15/30 correct in both control and treatment (0.0 pp observed difference; paired 95% bootstrap interval -16.7 to +20.0 pp). Jev was called on 20/30 treatment attempts, but 0/20 successful calls met the frozen context-completeness check. The resulting v2 manager treatment moves authoritative request construction into deterministic host code; its live qualification passed on an already-observed task and a fresh 30-pair v2 effectiveness cohort is preregistered but unrun.
The major changes that produced the current approach, with the deeper chronology and claim corrections available on demand.
February 2026
A Claude-to-Codex wrapper explored bounded work, observed evidence, and the difference between process exit and an agent’s claims.
invoke-codex-from-claudeArchivedMay 2026
An executor-neutral contract separated task, scope, outputs, and provenance. Copilot/manual packages and a mock executor were explored; the multi-executor runtime remained unfinished.
agent-jobUnfinished precursorJuly–August 2026
Delegation expands into isolated execution, restart lineage, durable evidence, independent review, and explicit acceptance, with a headless core.
Agent-WorkflowContinuing developmentReliability investigation
Structured intake and prompt delivery · archived
Operator console for PromptForge · archived
Prompt review and task coordination
Related investigations, not required stages or runtime dependencies.
26–27 Aug
Reduce routine workflow ceremony while preserving lifecycle boundaries.
30 Aug
SpecGen makes versioned decisions and exact source-baseline provenance part of the execution handoff.
September
Codebase Memory CLI changes the supported integration boundary for restricted environments, retaining the upstream engine and attribution.
18–19 Sep
Typed semantic evidence and neutral comparison semantics become independently inspectable.
23 Sep
Worker closeout simplification is paired with retained software and mixed evidence in benchmarking run 5.
24–26 Sep
A qualified DeepSeek-Flash adjudicator cohort produced the blinded frozen oracle first. The next gate was instrumentation rather than conclusions: an 8-case development-only TypeSafe/Jev smoke then verified 8 provider requests, 24 seam observations, 0 exclusions, oracle separation, probability persistence, privacy boundaries, and request-level accounting before the 120-case study.
26–27 Sep
Post-freeze review found that the authoritative v1 A/B/C pass retained labels but not structured per-decision justifications. A later audit found provider reasoning summaries in private Inspect logs, correcting the diagnosis from total loss to a promotion failure: useful explanatory evidence existed upstream but never entered the durable pass or original human-resolution workflow. The frozen oracle remains unchanged.
27 Sep
The full 120-case study completed with 360 observations, 120 provider requests, and 0 exclusions. Reporting preserved raw evidence when a 0.99 probability-mass edge case surfaced, and P3 review caught that the first publisher would have copied private oracle provenance. A redacted public projection then passed allowlist, manifest, sample-identity, and privacy verification before publication.
27 Sep
Building SpecGen and a versioned shared handoff made the planning/execution boundary explicit. Re-evaluating OpenSpec then showed that the planning side could standardize upstream while Agent-Workflow keeps source-agnostic execution and evidence contracts native. This is a target direction; qualification and migration remain pending.
1 Oct
A Jev skill that activated in controlled fixtures made zero calls across six real manager decisions. Agent-Workflow therefore prototypes a provider-neutral checkpoint outside the agent: draft first, neutral evidence projection second, explicit reconciliation on disagreement or uncertainty. The mechanism is merged; effectiveness remains unmeasured.
2 Oct
Thirty paired SWE-Lancer manager tasks produced 15/30 correct in both control and treatment: 0.0 pp observed difference, with a paired 95% bootstrap interval from -16.7 to +20.0 pp. Jev was called in 20 treatment attempts, but none of those calls met the frozen context-completeness check, turning request fidelity into the next engineering problem.
2 Oct
The v2 manager treatment separates generic request mechanics from task-specific completeness policy, reloads exact task and proposal text from the pinned source, preserves request identity through dispatch, and rejects unchanged semantic revisions. Live qualification passed on an already-observed task; the fresh v2 effectiveness cohort is preregistered and remains unrun.
Inspectable claim changes
Recent public-safe notes, browsed horizontally. Select one to inspect the claim change without lengthening the page.
A request-fidelity audit moved authoritative task context out of agent-written summaries and into deterministic host code.
Typed request schemas, payload bounds, exact receipts, and explicit skill guidance were enough to make a successful semantic call a faithful representation of the underlying decision problem.
In the fresh SWE-Lancer manager study, 20 treatment attempts successfully called Jev and every dispatched request was preserved exactly by the host-side receipt path.
None of those 20 calls passed the frozen context-completeness check. The complete authoritative task and proposal material fit inside the host limits; the loss occurred earlier when the coding agent summarized the source material before dispatch.
The next treatment generation uses a deterministic generic request builder plus a task-specific SWE-Lancer policy adapter. Exact title, description, and full proposal text are loaded from the pinned source; the agent contributes separately sourced repository evidence rather than recreating authoritative task content.
The v2 treatment has passed live qualification on an already-observed manager task and is preregistered for a fresh 30-pair cohort. That qualification proves request fidelity for the tested path, not an effectiveness improvement.
Independently usable applications, alongside the engineering tools and investigations above.
Grounded converts a natural-language analytics question into a structured QueryPlan using an LLM, validates the plan against allowlists, compiles parameterized SQL from approved fragments, executes it, and synthesizes an answer grounded in the returned rows.
Access-controlled demo: email verification and a lightweight access prompt protect metered AI resources from anonymous abuse.
Rocket Tool Sketcher is a browser-based designer for spindle and rammer tooling. It turns tube size and manufacturing tolerances into an annotated drawing and downloadable review/CAD artifacts.