SciencesLoop Open to AI Engineer roles
👋 Hi, I'm Lu.

I build grounded, testable AI agents for science & operations.

My work combines retrieval, evaluation, safety guardrails, and domain expertise — turning complex, high-context knowledge workflows into reliable software. I like agents that leave traces, demos you can actually test, and answers that say where they came from.

context engineering · loop engineering · harness engineering · safety gating

Before AI, 15+ yrs in energy science at Argonne / JCESR ·97 papers ·28 patents ·h-index 39 ·research archive →
01 — Case studies

Two systems I own end-to-end

One enterprise delivery, one independent build — the clearest proof of how I work.

● Enterprise · work case study

Academy-Wide Virtual Assistant

A member-facing agentic assistant for a ~38,000-member association

An MCP-tool agent grounded in a live enterprise CRM: it answers member questions from authoritative records and files service tickets when it can't. I owned the MCP tool layer, the read/write data path, and a multi-environment Azure deployment (Container Apps, Key Vault, managed identity, SSO).
MCP toolsDynamics CRMAzureSSO / OIDC
● Solo build · daily driver

“Second Brain” — Local-First Personal AI

A local-first personal AI that compiles my work stream into knowledge

Local RAG (ChromaDB + sentence-transformers behind Flask), a real-time meeting assistant (streaming Whisper + screenshot OCR + auto-synthesis), and a self-compiling Obsidian wiki — all on-device. Includes root-causing a native-library concurrency segfault and killing metered-API cost.
RAGstreaming STTvision / OCRlocal-first
02 — Featured AI systems

A few systems I've built

Each answers the same questions: what is it, who's it for, can you try it, and what did I actually build? See the two lines I build along →

● Live demo

ChemGraph Loop

A real agentic chemistry workflow you can run live

Ask a plain-language question about a small molecule and watch a real ChemGraph agent resolve a molecule, run a simulation, draw a 3D structure, and audit its result against a workflow guard.
agentic workflowLangGraph + ASEtool useworkflow guard
● Live demo

Condition-Monitoring Agent

A deterministic predictive-maintenance pipeline you can run live

A deterministic 10-stage remaining-useful-life analysis on public NASA turbofan data, with human decision gates, evidence-backed decision cards, and cited diagnostics.
deterministic pipelinepredictive maintenancehuman-in-the-loopauditable trace
● Live demo

Preventive Health Model Lab

A QLoRA fine-tuning study you can inspect end-to-end

A controlled QLoRA study comparing Gemma 3 and MedGemma on the same preventive-health task, synthetic data, and evaluation harness.
QLoRAMedGemma / Gemma 3eval harnesssynthetic data
● Live demo

Redox RFB Predictor

A public redox-potential demo that checks molecular identity before prediction

Enter a SMILES or chemical name, inspect PubChem-backed name options and an editable RDKit structure, then confirm it before the fixed Random Forest predicts on RedDB's DFT scale.
redox potentialRDKitPubChemhuman confirmationflow batteries
● Live demo

Molecular Discovery Workflow

Resolve identity, inspect the structure, then show what evidence is missing

A public, bounded version of a private molecular-discovery prototype: PubChem-backed identity resolution, human structure confirmation, RDKit descriptors, and explicit evidence gaps instead of invented property values.
PubChemRDKitevidence gapshuman confirmation
● Live demo

Guideline-Faithfulness Chat

Reports clinical-guideline positions while preserving each one's official strength

Ask a knee-osteoarthritis question; the answer reports what NICE NG226 and ACR/AF say, re-reading each source so the recommendation strength stays faithful instead of collapsing into one confident answer. A research demo of guideline faithfulness — not medical advice.
clinical guidelinesevidence verificationreliabilityno medical advice
● Public repository

Scientific Agent Lab

An evaluation & reproducibility layer for scientific AI

Not another science assistant — the layer below that asks “when can you trust the result?” A deterministic, zero-dependency pipeline that separates evidence from assumptions, flags missing required data, refuses to conclude on gaps, and ships a reproducible trace an eval harness scores.
Pythonevaluation harnessreproducibilitydeterministic · no LLM
● Private system overview

Property OS

Safety-gated real-time voice & SMS operations agent

Operations workflows where privacy, safety, and data boundaries matter as much as model capability: natural voice, a hard-coded safety screen, and strict data governance.
realtime voiceTwilio relayLLM safety screendata governance
03 — How I build

How I build reliable agents

Four engineering surfaces — how context, control, evaluation, and safety fit into one system.

→ input · grounding

Context Engineering

What goes into the model: grounded retrieval, citations, and attribution guardrails, so answers stay tied to real evidence.

→ control · the loop

Loop Engineering

The control loop — retrieve → reason → act → verify — built so behavior is observable and correctable, not a black box.

→ proof · evaluation

Harness Engineering

How I know it works: replayable eval harnesses, deterministic tests, and regression checks that make behavior measurable.

→ guardrail · safety

Safety Gating

The guardrails: hard-coded safety screens, scoped tools, privacy boundaries, and human review before anything ships.

Scientific foundation

My AI work is shaped by 15+ years in electrochemical energy research at Argonne & JCESR — flow batteries, electrolytes, and autonomous materials discovery. It's why I care about evidence, reproducibility, and safety boundaries.

View research archive →
05 — Contact

Let's build something that holds up.

Open to AI / agent engineering roles where rigor matters. Ask Loopi about my work, or say hi.