Skip to content
Nate G. / Vancouver, WA · US

Professional profile

Nathan Gallo

Agentic Systems / AI Platform Engineer

I build reliable autonomous execution systems with explicit boundaries around authority, evidence, evaluation, recovery, and accountability.

Built on more than a decade of production backend and systems engineering.

Vancouver, Washington · US-based · Open to senior US-based Agentic Systems / AI Platform engineering opportunities.

Software / systems / agents

Software Engineering
Work Atlas

Recent work on agent-assisted software engineering: making context, specifications, execution, judgment, and evaluation explicit enough to inspect and improve.

The Work Atlas records the systems, studies, failures, corrections, and evidence behind that work.

Engineering foundation

More than a decade of production software engineering across backend services, APIs, databases, distributed systems, cloud and CI/CD, greenfield development, brownfield systems, and legacy modernization. That background is the baseline for the recent agentic work below.

.NET / C#production backend

Agentic Engineering Ecosystem

Independent tools connected through specifications, contracts, orchestration, evidence, and evaluation.

Select a project to inspect its boundaries and evidence.

Context & repository intelligence

Source, structure, and codebase context.

Specification & handoff

Requirements, decisions, and durable baselines.

Execution & lifecycle

Isolation, recovery, review, and acceptance.

Semantic assurance

Bounded judgment with explicit uncertainty.

Measurement & evaluation

Comparison, produced software, and inspectable evidence.

Capabilities are facets. A project can appear in more than one; a shared concern does not imply a dependency.

No project selected

Engineering Questions

Each project began with a concrete engineering problem. The project, decision, and evidence links show how the answer changed as the work matured.

Engineering questionProject / workDesign responseInspect
Engineering questionCorporate environments may prohibit third-party MCP servers, but the need for repository indexing and codebase context remains.
Project / work

Codebase Memory MCP / CLI

CLI is my modified fork of DeusData/codebase-memory-mcp

Design responseOffer DeusData's upstream MCP or my modified local CLI fork according to environment policy.
Engineering questionRequirements and implementation decisions were being trapped in transient chat context and needed durable, versioned handoff artifacts.
Project / work

SpecGen-AW

Design responseMake requirements, decisions, and source baselines explicit and versioned before execution.
Engineering questionA worker process can finish without establishing that the work is complete, reviewed, or accepted.
Project / work

Agent-Workflow

Design responseSeparate execution from evidence, review, and deterministic acceptance; preserve enough state to recover and resume.
Engineering questionSome routing decisions benefit from semantic judgment, but lifecycle and policy authority still need to remain deterministic and inspectable.
Project / work

TypeSafe / Jev

Design responseBound semantic calls to typed evidence and explicit fallback; keep application policy responsible for the applied result.
Engineering questionAdditional orchestration is only defensible if its cost and effect can be measured against a simpler control.
Project / work

Benchmark tooling

Design responseRun paired experiments, preserve produced software and measurements, and expose limitations with the results.
Engineering questionA semantic system can disagree with deterministic routing without being more correct. How do we distinguish difference from improvement?
Project / work

Comparative Evaluation

Design responseFreeze public-safe cases and an independent blinded oracle before candidate/control comparison; use independent A/B labels, C only for disagreements, recorded resolution for genuine three-way conflicts, preserve enough decision provenance to diagnose disagreement, and evaluate raw semantic correctness separately from host policy.

Decisions & Evidence

Current principles and evidence records stay inspectable without turning the page into a vertical document wall.

What the work clarified

Completion needs an acceptance boundary

A worker finishing does not establish that the requested work is complete. Execution, evidence, review, and acceptance need distinct states, with explicit rules for accepting the result.

Completion gates
More clarified principles6 additional records

The common path should not carry unnecessary ceremony

Recovery and review mechanisms can add overhead to routine work. Preserve the boundaries that matter, simplify the common path, and measure the consequences.

Simplification decisions

A private evidence artifact is not automatically publishable

P3 review caught that the original publisher would have copied private oracle provenance. The published result therefore uses a separately defined public projection with explicit allowlist, integrity, sample-identity, and privacy verification.

Publication integrity

Multiple agents are not automatically independent evidence

A Codex-LB shutdown investigation deliberately separated code reading, empirical reproduction, an independent verifier using a different method, and orchestrator remeasurement. Later upstream review found an additional bounded-wait risk and replaced a merely correlated regression trigger with a deterministic barrier. The upstream PR remains under review.

Upstream Codex-LB review

Dogfooding is useful when frustration becomes an invariant

Recent Agent-Workflow use exposed four narrow contract failures: stale closeout guidance, equivalent cwd evidence scored as contradictory, recursive scope expressed as an exact path, and diagnostic loss when the configured state root was unavailable. Each became a bounded merged correction without rewriting the historical failure into a broader claim.

Diagnostic preservation correction

Semantic context needs a contract

More context can make a semantic answer more decisive without making it more discriminating. Primary artifacts and deterministic tool output should be separated from agent-authored claims; prior semantic answers and the choosing agent’s verdict should stay out of the independent provider projection by default.

Neutral projection rules

Request validity is not evidence completeness

A bounded, typed, successfully dispatched semantic request can still omit the authoritative evidence needed for the decision. In the first fresh manager study, 0/20 successful Jev calls met the frozen completeness check even though the bridge preserved every dispatched string. Request construction therefore became a deterministic, versioned part of the treatment.

SWE-manager v2 request policy

Inspectable evidence records

Evidence in context

TypeSafe / Jev

Jevtyped evidence

Typed semantic judgment returns bounded evidence while application code retains policy, side effects, lifecycle, review, and acceptance authority.

Inspect evidence boundary and integration detailTyped receipts · qualification evidence

The integration explores bounded semantic judgment through typed Choice, Noul, and Score evidence. Receipts preserve provenance: source references, question versions, model details, request hashes, and outcome status. Uncertainty, no-decision outcomes, and fallback remain explicit.

Application code determines how evidence affects the result. TypeSafe/Jev does not own orchestration, side effects, review, or acceptance, and can be used independently of Agent-Workflow.

Structural checkpoint update. Controlled skill qualification did not translate into reliable optional uptake: six frozen real manager decisions made zero Jev calls. Agent-Workflow therefore prototypes a provider-neutral outer checkpoint that records the choosing agent’s draft before semantic evidence, projects neutral primary evidence without that verdict, and requires explicit reconciliation on disagreement or uncertainty.

Context discipline. Later review work reinforced that richer context can make semantic output more decisive without making it more discriminating. The independent projection now separates deterministic tool evidence from agent-authored claims and excludes the choosing agent’s preferred answer and prior Jev answers by default.

Fresh manager effectiveness study. A 30-pair SWE-Lancer manager study then observed 15/30 correct in both control and treatment (0.0 pp difference; paired 95% bootstrap interval -16.7 to +20.0 pp). Jev was invoked on 20 treatment attempts, but 0/20 successful calls met the frozen context-completeness check. The result is preserved as evidence for the treatment that actually ran rather than retroactively repaired.

Request-fidelity correction. The audit showed that the bridge did not truncate the successful calls; the coding agent had compressed authoritative task and proposal material before dispatch. The next treatment generation therefore builds the neutral request in deterministic host code. That v2 path has passed live qualification on an already-observed task; its fresh effectiveness cohort is preregistered but unrun.

Evidence boundary. Provider success alone does not establish correctness. The completed routing-semantic-v1 study now supplies independent seam-level correctness evidence for a frozen 120-case cohort; those findings should not be generalized into universal Jev superiority or downstream software-quality causality.

benchmarking run 5 — one development pair

evaluationbenchmark + evidence

One retained development pair used approximately 185 seconds less wall time and 235,308 fewer provider tokens with the slimmed Agent-Workflow treatment, while scoring 2.5 machine-score points lower.

Boundary: one development pair, no human-complete pair, and no measured cost saving.

Inspect retained artifacts and limitationsControl + treatment evidence

Paired benchmark runs retain the produced software alongside timing, token usage, scoring, and methodology. This makes changes to the process inspectable against an explicit control.

This was one development pair and is not a generalized performance claim.

The pair is awaiting human review; no human-complete pair is recorded. Cost evidence is unavailable, so fewer tokens should not be presented as a measured cost saving.

TypeSafe qualification ran before both treatments, with semantic routing outside the comparison. These results do not establish a TypeSafe effect.

Comparative Evaluation

evaluationindependent correctness

The study moved from asking whether semantic routing behaves differently to asking whether it is actually more correct against independently frozen ground truth.

Task class81.67% semantic vs 52.50% control
Interaction required91.67% semantic vs 71.67% control
Semantic riskMAE 0.28975 vs 0.29167 · interval spans zero

Current boundary: v1 is published. V2 achieved a full 0.6.3 qualification pass, but the 0.6.4 model-facing schema representation changed and must qualify again before real A/B. No real v2 A/B/C result is claimed.

Inspect methodology, corrections, and study artifactsOracle · provenance · publication

The methodological shift is the important part of this work. Early instrumentation could show that TypeSafe/Jev and deterministic routing sometimes behaved differently. That was not enough to say which decision was better.

Earlier questionDoes TypeSafe/Jev behave differently?
Better questionWhen it behaves differently, is it actually more correct?
NeededIndependent ground truth
ThereforeFrozen cases + blinded adjudication + a frozen oracle before comparison

Independent ground truth is now frozen. The qualified DeepSeek-Flash adjudicator cohort completed independent A and B labeling on the same hash-pinned blinded view, C received only disputed cases and seams, and genuine three-way conflicts went through recorded human resolution. The resulting oracle was frozen and validated before any live TypeSafe/Jev comparative inference.

Oracle construction now has a recorded cost/usage baseline. The initial qualification and adjudication runs through P0B oracle freeze used the OpenAI Codex harness, routed through codex-lb to the DeepSeek-4.1-Flash API. Provider accounting for that baseline was $0.07 USD across 36 API requests and 726,263 tokens. These totals stop at oracle freeze and do not include P1/P2 Jev traffic.

The staged execution path is now complete. P1 first verified the live evidence path on 8 development-only cases. P2 then ran all 120 frozen cases without exposing the oracle during inference, producing 360 observations from 120 provider requests with 0 exclusions. P3 independently verified the sanitized publication projection before the result was published.

The result is deliberately heterogeneous. In this frozen cohort, the semantic candidate was more accurate on task-class and interaction-required classification, while semantic-risk ordinal error was approximately unchanged at the study's resolution. All 120 provider requests succeeded. Provider treatment-cost evidence is incomplete.

Raw semantic quality and operational policy remain separate layers. Choice, Noul, and Score evidence are evaluated against the oracle without treating the current confidence threshold, fallback rule, or non-automatable risk policy as model correctness.

Independent adjudicationLabels + intended decision justification
Persisted v1 recordLabels retained
Authoritative v1 passLabels retained; structured justifications omitted
Next cohort contractStructured justification + round-trip persistence check

Decision-provenance correction. The authoritative v1 A/B/C passes retained labels but no structured per-decision justifications, so the original human resolver worked from the case, metadata, rubric, and labels. A later audit found contemporaneous provider reasoning summaries in private Inspect logs. Those summaries were supplementary execution evidence, were not available to the resolver, are not hidden chain-of-thought, and are not being retrofitted into the frozen oracle. The failure was promotion of useful evidence into the durable decision record, not complete historical loss.

Reporting correction. P2 exposed one 0.99-total probability vector in each multiclass seam. Raw evidence was preserved; only the derived calibration copy was normalized, and the deviation was disclosed.

Publication correction. P3 review found that the first publisher would have copied the private oracle verbatim. The released bundle instead uses a redacted public oracle projection and passed an independent allowlist, manifest-integrity, cohort-identity, and privacy verifier before publication.

Results boundary. Routing Semantic v1 is complete and P3-published: 120 frozen cases produced 360 decision observations from 120 semantic provider requests with 0 exclusions. In this cohort, semantic-candidate accuracy was 81.67% vs 52.50% for task class and 91.67% vs 71.67% for interaction required; both paired 95% accuracy-difference intervals remained above zero. Semantic-risk MAE was 0.28975 vs 0.29167 and its paired interval spanned zero. All 120 provider requests succeeded. Provider treatment-cost evidence is incomplete, so no cost-advantage, universal-superiority, or downstream software-quality claim is supported. Routing-semantic-v2 later reached a first full live IA-1 through IA-11 qualification pass under Benchmark 0.6.3 with sanitized proof that strict JSON-schema controls reached Codex-LB ingress. Before any real v2 label was produced, real-cohort readiness review found that the 120-case model-facing schema expanded to 225,971 serialized characters because an identical record shape was repeated per case. Benchmark 0.6.4 replaces that representation with the eligibility-grouped-case-enum/v1 strategy, reducing the equivalent schema to 3,310 characters while preserving all 120 case identities. Because the model-facing schema changed, the earlier pass cannot authorize the 0.6.4 cohort: a fresh live IA-1 through IA-11 qualification is required. Real routing-semantic-v2 A/B/C remains unrun. Separately, a fresh 30-pair Agentic-Jev SWE-manager v1 study completed with 15/30 correct in both control and treatment (0.0 pp observed difference; paired 95% bootstrap interval -16.7 to +20.0 pp). Jev was called on 20/30 treatment attempts, but 0/20 successful calls met the frozen context-completeness check. The resulting v2 manager treatment moves authoritative request construction into deterministic host code; its live qualification passed on an already-observed task and a fresh 30-pair v2 effectiveness cohort is preregistered but unrun.

Study state
P0A–P3 complete · sanitized public result published
Full cohort
120 cases · 360 observations · 120 provider requests · 360 oracle outcomes · 0 exclusions
Classification seams
Task class +29.17 pp · interaction required +20.0 pp · both paired 95% intervals above zero
Semantic risk
MAE 0.28975 semantic vs 0.29167 control · paired interval spans zero
Request evidence
120/120 success · 91,445 input tokens · 10,598 output · 102,043 total · p95 0.2206 s
Decision provenance
Structured justifications absent from authoritative v1 pass · private reasoning summaries retained upstream only
Publication
Public oracle projection · exact-file allowlist · manifest integrity · privacy checks PASS
Claim boundary
No universal superiority, downstream quality causality, treatment-cost advantage, or material semantic-risk improvement claim

Evolution / Field Notes

The major changes that produced the current approach, with the deeper chronology and claim corrections available on demand.

Major inflection points 2026

  1. February 2026

    Delegation becomes explicit

    A Claude-to-Codex wrapper explored bounded work, observed evidence, and the difference between process exit and an agent’s claims.

    invoke-codex-from-claudeArchived
  2. May 2026

    The job becomes a contract

    An executor-neutral contract separated task, scope, outputs, and provenance. Copilot/manual packages and a mock executor were explored; the multi-executor runtime remained unfinished.

    agent-jobUnfinished precursor
  3. July–August 2026

    Work gains a durable lifecycle

    Delegation expands into isolated execution, restart lineage, durable evidence, independent review, and explicit acceptance, with a headless core.

    Agent-WorkflowContinuing development
Research branches Intake, coordination, and reliability

Related investigations, not required stages or runtime dependencies.

Later ecosystem development12 recorded milestones · August–October 2026
  1. 26–27 Aug

    Simplify the common path

    Reduce routine workflow ceremony while preserving lifecycle boundaries.

    Inspect change
  2. 30 Aug

    Bind the handoff to source

    SpecGen makes versioned decisions and exact source-baseline provenance part of the execution handoff.

    Inspect change
  3. September

    Move repository context to a CLI

    Codebase Memory CLI changes the supported integration boundary for restricted environments, retaining the upstream engine and attribution.

    Inspect change
  4. 18–19 Sep

    Separate judgment from authority

    Typed semantic evidence and neutral comparison semantics become independently inspectable.

    Inspect change
  5. 23 Sep

    Measure the simplification

    Worker closeout simplification is paired with retained software and mixed evidence in benchmarking run 5.

    Inspect change
  6. 24–26 Sep

    Freeze ground truth, then verify the evidence path

    A qualified DeepSeek-Flash adjudicator cohort produced the blinded frozen oracle first. The next gate was instrumentation rather than conclusions: an 8-case development-only TypeSafe/Jev smoke then verified 8 provider requests, 24 seam observations, 0 exclusions, oracle separation, probability persistence, privacy boundaries, and request-level accounting before the 120-case study.

    Inspect change
  7. 26–27 Sep

    Treat decision provenance as evidence

    Post-freeze review found that the authoritative v1 A/B/C pass retained labels but not structured per-decision justifications. A later audit found provider reasoning summaries in private Inspect logs, correcting the diagnosis from total loss to a promotion failure: useful explanatory evidence existed upstream but never entered the durable pass or original human-resolution workflow. The frozen oracle remains unchanged.

    Inspect change
  8. 27 Sep

    Publish the result through a fail-closed boundary

    The full 120-case study completed with 360 observations, 120 provider requests, and 0 exclusions. Reporting preserved raw evidence when a 0.99 probability-mass edge case surfaced, and P3 review caught that the first publisher would have copied private oracle provenance. A redacted public projection then passed allowlist, manifest, sample-identity, and privacy verification before publication.

    Inspect change
  9. 27 Sep

    Discover the upstream standardization boundary

    Building SpecGen and a versioned shared handoff made the planning/execution boundary explicit. Re-evaluating OpenSpec then showed that the planning side could standardize upstream while Agent-Workflow keeps source-agnostic execution and evidence contracts native. This is a target direction; qualification and migration remain pending.

    Inspect change
  10. 1 Oct

    Move semantic checking outside the choosing agent

    A Jev skill that activated in controlled fixtures made zero calls across six real manager decisions. Agent-Workflow therefore prototypes a provider-neutral checkpoint outside the agent: draft first, neutral evidence projection second, explicit reconciliation on disagreement or uncertainty. The mechanism is merged; effectiveness remains unmeasured.

    Inspect change
  11. 2 Oct

    Measure the first fresh manager effectiveness cohort

    Thirty paired SWE-Lancer manager tasks produced 15/30 correct in both control and treatment: 0.0 pp observed difference, with a paired 95% bootstrap interval from -16.7 to +20.0 pp. Jev was called in 20 treatment attempts, but none of those calls met the frozen context-completeness check, turning request fidelity into the next engineering problem.

    Inspect change
  12. 2 Oct

    Move authoritative semantic context into deterministic code

    The v2 manager treatment separates generic request mechanics from task-specific completeness policy, reloads exact task and proposal text from the pinned source, preserves request identity through dispatch, and rejects unchanged semantic revisions. Live qualification passed on an already-observed task; the fresh v2 effectiveness cohort is preregistered and remains unrun.

    Inspect change

Inspectable claim changes

Field Log

Recent public-safe notes, browsed horizontally. Select one to inspect the claim change without lengthening the page.

View all 10 field notes
01 / 07

Method Change

Request construction became part of the treatment

A request-fidelity audit moved authoritative task context out of agent-written summaries and into deterministic host code.

Agentic-Jev · request fidelityv2 live-qualified · fresh effectiveness cohort preregistered, not run
Belief

Typed request schemas, payload bounds, exact receipts, and explicit skill guidance were enough to make a successful semantic call a faithful representation of the underlying decision problem.

Evidence

In the fresh SWE-Lancer manager study, 20 treatment attempts successfully called Jev and every dispatched request was preserved exactly by the host-side receipt path.

Contradiction

None of those 20 calls passed the frozen context-completeness check. The complete authoritative task and proposal material fit inside the host limits; the loss occurred earlier when the coding agent summarized the source material before dispatch.

Correction

The next treatment generation uses a deterministic generic request builder plus a task-specific SWE-Lancer policy adapter. Exact title, description, and full proposal text are loaded from the pinned source; the agent contributes separately sourced repository evidence rather than recreating authoritative task content.

Current claim boundary

The v2 treatment has passed live qualification on an already-observed manager task and is preregistered for a fresh 30-pair cohort. That qualification proves request fidelity for the tested path, not an effectiveness improvement.

Live Hosted Applications

Independently usable applications, alongside the engineering tools and investigations above.

Grounded

Grounded converts a natural-language analytics question into a structured QueryPlan using an LLM, validates the plan against allowlists, compiles parameterized SQL from approved fragments, executes it, and synthesizes an answer grounded in the returned rows.

Access-controlled demo: email verification and a lightweight access prompt protect metered AI resources from anonymous abuse.

Rocket Tool Sketcher

Rocket Tool Sketcher is a browser-based designer for spindle and rammer tooling. It turns tube size and manufacturing tolerances into an annotated drawing and downloadable review/CAD artifacts.