ds.Let’s connect
DANIEL SHDEED / FULL STACK ENGINEER

Build the system.
Question the intelligence.

I build full-stack systems and study what makes AI reliable. My work connects applied agents, rigorous evaluation, and the mechanics of how models are stored, learn, and forget.

THE WORKING LOOP01—04
evidenceat the center
01 / BUILD02 / MEASURE03 / QUESTION04 / REPEAT
Systems thinkingScientific curiosity
A practice in engineering & independent researchAPPLIED AI / EVALUATION / CONTINUAL LEARNING
01 / SELECTED WORK

Selected work.
One line of inquiry.

How do we build intelligent systems whose actions, explanations, and results we can actually trust?

BOUNDED AUTONOMY / HUMAN CONTROL
01Investigate3 typed read tools
02GroundValidate evidence IDs
03ApproveHuman decision required
3 turns6 tool calls90-second bound
01 / APPLIED AILocal implementation

OpsPilot

An agent with a chain of accountability.

Investigate shipment exceptions, retrieve relevant policies, and propose an action—with evidence attached and a supervisor in control.

FastAPIReactPostgres · pgvectorCelery
GitHub
EXACT EVALUATION / NO LLM JUDGES
score =100 × vv*
v achieved changev* certified optimum
4 × 4 × 4 world≤ 3 relocations3 fixed shadows
02 / LLM EVALUATIONPilot pending

Shadow Twins

Same shadows. Different structure.

A constrained spatial benchmark: rearrange a voxel solid without changing its projections. Score every answer against an exhaustively certified optimum.

PythonReact · Three.jsExact solver
CONTROLLED EXPERIMENT / PAIRED SEEDS
Shared training prefixweb text
Continue webcontrol
Switch to Pythondomain shift
Compare held-out web deterioration
RMS vs. TaperNorm3 paired seeds
03 / INDEPENDENT RESEARCHPreparation stage

Domain-Shift Forgetting

What does a model leave behind?

A controlled pilot studying whether internal normalization changes persistent forgetting when language-model training shifts from web text to Python.

Transformer trainingPaired experimentsReproducibility
GitHub
TIME-AWARE ANALYTICS / GROUNDED EXPLANATIONS
Race clockt
Your broadcastt − delay
Evidence time ≤ viewer time
One event pipelineLive + replay
04 / REAL-TIME SYSTEMSBeta

FastPitStop

Race intelligence. Without the spoilers.

A second-screen race companion combining deterministic analytics, confidence-aware predictions, and cited AI explanations—all synchronized to the viewer’s clock.

SvelteKitPython · PolarsTimescaleDB
GitHub
MODEL INTERNALS / BUILT FROM RAW BYTES
$ biopsy inspect model.gguf
$ biopsy stats model.gguf
$ biopsy health model.gguf
$ biopsy diff base tuned
safetensors + GGUFScalar + AVX2
05 / AI SYSTEMS TOOLINGNear completion

biopsy

Understand the model beneath the API.

A Rust CLI that inspects neural-network files from raw bytes. Hand-written parsers, streaming tensor statistics, weight health checks, and checkpoint comparisons make model internals tangible.

RustQuantizationMemory mappingSIMD · Rayon
GitHub
THE THREAD CONNECTING IT ALL

Intelligence is useful.
Accountability makes it usable.

01

Evidence before explanation

A fluent answer is only a starting point. Claims should connect to sources, constraints, and the information available at the time.

02

Evaluation before confidence

Use exact checks where possible. Keep baselines, uncertainty, failure categories, and unfinished validation visible.

03

Control before autonomy

Bound the agent. Make approval explicit. Design for retries, interrupted workers, and uncertain outcomes.

02 / THE RESEARCH AGENDA

The next good question.

Directions I want to explore next. These are proposals, not completed projects or reported findings.

R.01
AGENT RELIABILITY

When should an agent abstain?

Measure the tradeoff between useful action and unsupported confidence.

Proposed

The experiment

Extend OpsPilot with controlled missing, stale, and contradictory evidence. Compare always-answer, rule-based abstention, and confidence-threshold policies on the same reviewed cases.

What would count as progress

Report unsupported-claim rate, useful completion rate, escalation frequency, and latency across evidence conditions. Publish the cases and failure taxonomy alongside the curves.

R.02
REPRESENTATION & REASONING

Same problem. Different coordinates.

Does a model’s spatial reasoning survive a change in representation?

Proposed

The experiment

Build on Shadow Twins’ planned symmetry comparisons. Rotate equivalent objects, permute coordinate descriptions, and compare answers while holding the certified optimum fixed.

What would count as progress

Measure paired score differences and invalid-answer categories. Keep variants clustered with the original instance; distinguish representation sensitivity from task difficulty.

R.03
CONTINUAL LEARNING

Can forgetting be caught early?

Explore whether early diagnostics anticipate persistent deterioration.

Proposed

The experiment

After the fixed Stage 1 pilot, design a separate exploratory study of checkpoint-level loss and representation diagnostics around a domain switch. Keep the existing pilot protocol unchanged.

What would count as progress

Test prespecified early-warning signals on held-out runs. Compare against simple loss-based baselines and report false alarms as well as detection lead time.

03 / THE PERSON & THE PRACTICE

Engineer by practice.
Researcher by curiosity.

I’m Daniel Shdeed, a Full Stack Engineer building toward AI engineering through hands-on systems work and independent study.

I’m interested in the space between a model that produces an answer and a system that deserves to be trusted. That means understanding the data path, the failure modes, and the experiment—not just the API.

Find me on LinkedIn
MY SELF-STUDY MAP Learning through building
01

From interface to infrastructure

Typed APIs, data models, asynchronous workers, replayable event streams, and observable execution.

Practiced through OpsPilot & FastPitStop
02

From model output to evidence

Retrieval, citation validation, constrained evaluation, exact scoring, and uncertainty-aware reporting.

Practiced through OpsPilot & Shadow Twins
03

From implementation to experiment

Transformer internals, normalization, paired controls, leakage prevention, and reproducible checkpoints.

Studied through Domain-Shift Forgetting
04

From raw bytes to model internals

Binary formats, floating-point decoding, block quantization, streaming statistics, and measured SIMD optimization.

Practiced through biopsy in Rust
LET’S BUILD SOMETHING THAT HOLDS UP.

Good systems start with
good questions.

Connect on LinkedIn

Interested in AI engineering, evaluation, and research-minded teams.

PROJECT FIELD NOTES