OpsPilot
An agent with a chain of accountability.
Investigate shipment exceptions, retrieve relevant policies, and propose an action—with evidence attached and a supervisor in control.
I build full-stack systems and study what makes AI reliable. My work connects applied agents, rigorous evaluation, and the mechanics of how models are stored, learn, and forget.
How do we build intelligent systems whose actions, explanations, and results we can actually trust?
An agent with a chain of accountability.
Investigate shipment exceptions, retrieve relevant policies, and propose an action—with evidence attached and a supervisor in control.
Same shadows. Different structure.
A constrained spatial benchmark: rearrange a voxel solid without changing its projections. Score every answer against an exhaustively certified optimum.
What does a model leave behind?
A controlled pilot studying whether internal normalization changes persistent forgetting when language-model training shifts from web text to Python.
Race intelligence. Without the spoilers.
A second-screen race companion combining deterministic analytics, confidence-aware predictions, and cited AI explanations—all synchronized to the viewer’s clock.
Understand the model beneath the API.
A Rust CLI that inspects neural-network files from raw bytes. Hand-written parsers, streaming tensor statistics, weight health checks, and checkpoint comparisons make model internals tangible.
A fluent answer is only a starting point. Claims should connect to sources, constraints, and the information available at the time.
Use exact checks where possible. Keep baselines, uncertainty, failure categories, and unfinished validation visible.
Bound the agent. Make approval explicit. Design for retries, interrupted workers, and uncertain outcomes.
Directions I want to explore next. These are proposals, not completed projects or reported findings.
Measure the tradeoff between useful action and unsupported confidence.
Extend OpsPilot with controlled missing, stale, and contradictory evidence. Compare always-answer, rule-based abstention, and confidence-threshold policies on the same reviewed cases.
Report unsupported-claim rate, useful completion rate, escalation frequency, and latency across evidence conditions. Publish the cases and failure taxonomy alongside the curves.
Does a model’s spatial reasoning survive a change in representation?
Build on Shadow Twins’ planned symmetry comparisons. Rotate equivalent objects, permute coordinate descriptions, and compare answers while holding the certified optimum fixed.
Measure paired score differences and invalid-answer categories. Keep variants clustered with the original instance; distinguish representation sensitivity from task difficulty.
Explore whether early diagnostics anticipate persistent deterioration.
After the fixed Stage 1 pilot, design a separate exploratory study of checkpoint-level loss and representation diagnostics around a domain switch. Keep the existing pilot protocol unchanged.
Test prespecified early-warning signals on held-out runs. Compare against simple loss-based baselines and report false alarms as well as detection lead time.
I’m Daniel Shdeed, a Full Stack Engineer building toward AI engineering through hands-on systems work and independent study.
I’m interested in the space between a model that produces an answer and a system that deserves to be trusted. That means understanding the data path, the failure modes, and the experiment—not just the API.
Find me on LinkedInTyped APIs, data models, asynchronous workers, replayable event streams, and observable execution.
Practiced through OpsPilot & FastPitStopRetrieval, citation validation, constrained evaluation, exact scoring, and uncertainty-aware reporting.
Practiced through OpsPilot & Shadow TwinsTransformer internals, normalization, paired controls, leakage prevention, and reproducible checkpoints.
Studied through Domain-Shift ForgettingBinary formats, floating-point decoding, block quantization, streaming statistics, and measured SIMD optimization.
Practiced through biopsy in RustInterested in AI engineering, evaluation, and research-minded teams.