I build grounded, testable AI agents for science & operations.
My work combines retrieval, evaluation, safety guardrails, and domain expertise — turning complex, high-context knowledge workflows into reliable software. I like agents that leave traces, demos you can actually test, and answers that say where they came from.
context engineering · loop engineering · harness engineering · safety gating
Two systems I own end-to-end
One enterprise delivery, one independent build — the clearest proof of how I work.
Academy-Wide Virtual Assistant
A member-facing agentic assistant for a ~38,000-member association
“Second Brain” — Local-First Personal AI
A local-first personal AI that compiles my work stream into knowledge
A few systems I've built
Each answers the same questions: what is it, who's it for, can you try it, and what did I actually build? See the two lines I build along →
ChemGraph Loop
A real agentic chemistry workflow you can run live
Condition-Monitoring Agent
A deterministic predictive-maintenance pipeline you can run live
Preventive Health Model Lab
A QLoRA fine-tuning study you can inspect end-to-end
Redox RFB Predictor
A public redox-potential demo that checks molecular identity before prediction
Molecular Discovery Workflow
Resolve identity, inspect the structure, then show what evidence is missing
Guideline-Faithfulness Chat
Reports clinical-guideline positions while preserving each one's official strength
Scientific Agent Lab
An evaluation & reproducibility layer for scientific AI
Property OS
Safety-gated real-time voice & SMS operations agent
How I build reliable agents
Four engineering surfaces — how context, control, evaluation, and safety fit into one system.
Context Engineering
What goes into the model: grounded retrieval, citations, and attribution guardrails, so answers stay tied to real evidence.
Loop Engineering
The control loop — retrieve → reason → act → verify — built so behavior is observable and correctable, not a black box.
Harness Engineering
How I know it works: replayable eval harnesses, deterministic tests, and regression checks that make behavior measurable.
Safety Gating
The guardrails: hard-coded safety screens, scoped tools, privacy boundaries, and human review before anything ships.
Field notes
Working notes, build logs, and small lessons from building grounded agents. Some are polished. Some are still growing.
All notes →My AI work is shaped by 15+ years in electrochemical energy research at Argonne & JCESR — flow batteries, electrolytes, and autonomous materials discovery. It's why I care about evidence, reproducibility, and safety boundaries.
Let's build something that holds up.
Open to AI / agent engineering roles where rigor matters. Ask Loopi about my work, or say hi.