HELLO THERE · DATA ENGINEER, GÖTTINGEN, DE

hello there · same person, after hours

Ojaswi Sharma

I build data systems that prove themselves.

End-to-end pipelines, streams and simulations, instrumented until the numbers speak for the work. Previously: production healthcare data platforms for Fortune-500 pharma. Currently: MSc Applied Data Science, Göttingen.

By day the receipts do the talking. By night: artificial villagers, ideas on shelves, a dog named Kaju who lives too far away, and whatever's currently streaming through the workshop.

latest Fathom · P99 63.4 ms status 10-Bit Village live · 200K+ ticks focus data engineering · analytics · BI my uptime … yrs (the human, not the pipeline)

ORDER ISN’T CREATED. IT’S EXTRACTED.

EXPERIENCE

the humane version

SEP 2024 — MAR 2026 · 18 MONTHS

Veersa Technologies

DATA ENGINEER · AUG 2025 — MAR 2026

  • Owned 3–4 concurrent client pipelines end-to-end on Databricks + Delta Lake: medallion architecture, SCD Type 1 & 2, millions of records for Fortune-500 pharma / US healthcare
  • Agent-productivity BI data product: translated stakeholder KPI definitions into Databricks gold tables from 5+ sources (UKG, Pega, Siebel, Amazon Connect, SharePoint), powering Power BI/Excel reporting for daily, weekly, monthly, and quarterly productivity analysis
  • Pharmacy-fulfillment gold layer: multi-join fact tables across Snowflake + Databricks, watermark-based SCD ingestion, ACID Delta transactions
  • Diagnosed a silent data-loss incident: an upstream schema change had been dropping Bronze→Silver records undetected for months. Shipped a backward-compatible fix and backfilled all of history

DATA ENGINEERING INTERN · SEP 2024 — JUL 2025

  • Intern → Data Engineer in 11 months · Rising Star Q3 ’25
  • Joined at the start of the client’s migration off a legacy analytics stack onto Databricks for all their analytics needs, part of the founding team that built the platform from the ground up
  • Built a configuration-driven ingestion framework that cut new-client onboarding from 2 weeks to 3 days
  • Automated SQL-based data-quality checks with Databricks Alerts; orchestrated it all through Databricks Jobs

Where I learned to own things completely: the pipeline, the mistake, the fix, and the conversation with the C-suite about all three.

JUL — AUG 2024

Indika AI

AI/ML INTERN · JUL — AUG 2024

  • Fine-tuned OpenAI Whisper via Hugging Face Transformers for Indian-language ASR: 30–40% word-error-rate reduction
  • Lived the whole Hugging Face loop: open models fine-tuned on Indic-language datasets pulled straight from the Hub, inference squeezed into strictly limited GPU budgets
  • Built the training-data pipeline (automated audio extraction via the YouTube Data API) plus an internal API project feeding the same corpus
  • Helped write the team’s technical blog posts. First practice explaining models to people who don’t build them

Eight weeks, one clear number. Short chapters count when they end with evidence.

MEANWHILE, OFF THE CLOCK

  • IEEE MSIT · Executive Committee · 2023–25 · ran technical events, grew the branch
  • DevSource · Technical Content Lead · 2023–25 · led content for a dev community
  • Unnat Bharat Abhiyan · 2022 · sustainability & hygiene camp at a Delhi govt school

SELECTED WORK

receipts included · stories inside

Fathom. market data, live.

Real-time market data pipeline: Kafka streaming, L2 order-book reconstruction, sequence-gap detection, and a derived microstructure feed: microprice, order-book imbalance, effective spread. Built in four sealed phases, each with evidence.

P99 TICK-TO-STORE63.4 ms
PHASES SEALED 4 / 4
  • 01 INGEST · Binance live WS + CME replay, one path, Protobuf on the wire
  • 02 BOOK · full L2 reconstruction, sequence-gap detection
  • 03 DERIVE · microprice, imbalance, spread, published as a topic
  • 04 PROVE · 1h live soak, gap-free, P99 63.4 ms on a dashboard

sealed = proven with evidence, then closed. no phase reopened; each one shipped something the next one stood on.

SOAK1h live, gap-free
THE FULL STORY ↓

Why this project. Quant data systems live or die on market data infrastructure, and every tutorial pretends it’s easy. I wanted to see market data the way an exchange does: not in a notebook, but as a living stream with latency budgets, sequence numbers, and failure modes that only show up at 3am. This is the gap between “did a Kafka course” and “ran a pipeline through a live soak and can show you the P99.”

Why this stack. Kafka (Redpanda) for the backbone, Protobuf on the wire, TimescaleDB and ClickHouse at rest, Prometheus + Grafana watching everything. Boring choices, deliberately. The interesting part is what flows through them, and the discipline was proving each phase with evidence before sealing it.

What it achieves. Binance live WebSocket and Databento CME-futures replay running through the same path with parity. Full L2 order-book reconstruction with sequence-gap detection. A derived microstructure topic (microprice, order-book imbalance, effective spread) that other systems can subscribe to. FathomPulse is designed as the first subscriber. P99 of 63.4 ms tick-to-store, measured over a one-hour live soak, gap-free.

WHAT I DIDN'T KNOW GOING IN

  • day one: did not know what an L2 order book actually was
  • learned Protobuf because JSON was embarrassing me on the wire
  • found out 'exactly-once delivery' is a bedtime story vendors tell

phase 3 nearly broke me: sequence gaps silently corrupting the book for two days. the fix is the best code in the repo.

10-Bit Village. 100 tiny minds.

An emergent-society simulation: 100 agents, one world, no scripts. V2 asks the uncomfortable question: what do they remember, and does it matter?

V1 retired
  • shipped June 2026: deterministic C++ core, free-thought cognition, live HF Space
  • 3-day audit found it: 0.6% grounded actions, one myth on infinite loop, no persona in retellings
  • honest data, not a flattering demo: retired as a control dataset instead of quietly patched

a run that converges to a boring attractor is still evidence

V2 live now
  • 208,500+ ticks, running continuously since the Run 2 cutover
  • 51,186 gossips traded, 20,669 dreams, 33,127 grounded moments
  • 4,090 words coined and spreading through the population

snapshot taken 2026-08-03, the counter only goes up

RESIDENTSalive
THE FULL STORY ↓

Why this project. Because building pipelines teaches you how systems behave, but building a society teaches you how systems surprise you. This one is for fun, and it turned out to be the most serious thing I’ve made: 100 agents with 10 bits of mind each, dropped into one world with no scripts, developing trade, naming, and grudges on their own.

Why this stack. A deterministic simulation core (AVX2-pinned so every run is reproducible bit-for-bit), checkpointed engine state, agents deliberately small enough to reason about completely. When something emerges, I can prove it emerged. Determinism is what separates an experiment from a screensaver.

V1: what shipped. A two-segment cognition loop I call “free mind, hidden hands”: each tick, a villager gets one unconstrained thought, then a non-LLM C++ resolver reads that thought and extracts a valid action from it, with a thin survival reflex underneath so nobody starves to death by accident. Early versions caged the thought in a JSON grammar, not just the output, and the model went quiet the moment it was constrained. Freeing the thought and moving the constraints to the edges is what made it come alive. Shipped public on a Hugging Face Space: sprite-based live viewer, gossip, dreams, word-coinage, a replay tool.

Why V2. A 3-day audit of the live run said something a demo never would: the village had converged to a boring attractor by day one. Gossip drowned out grounded action (0.6% of events), retellings carried no trace of who was telling them, one myth looped on repeat because salience always picked the same “winner,” and most coined words were inflection noise, not real coinage. I retired that run as a control dataset instead of patching it quietly and calling it done.

What changed in V2. A grounded-action floor so villagers aren’t only gossiping, personality-conditioned prose feeding gossip and dreams instead of flavor-only text, a fatigue penalty so the same retelling can’t loop forever, sentence-aware trimming so distortion doesn’t sever mid-word, and a stricter coinage filter. Then a real cutover: archived Run 1’s full 77,926-tick history as a permanent dataset, wiped state, rebooted clean.

Honest read on quality. The writing is fluent and does invent real lore, place names and character names, unprompted: that’s the actual thesis paying off. What hasn’t fully landed: the trait-differentiated “distinct voice per personality” goal. Most villagers still read in a similar register right now. That’s the next thing to chase, not a claim I’m making today.

WHAT I DIDN'T KNOW GOING IN

  • started with zero idea how to make determinism survive a rebuild
  • learned that AVX2 changes float math. the hard way. twice
  • did not expect to feel guilty pausing a simulation

the villagers renamed the river three times before i stopped correcting them. that's when it got interesting.

"Nate," he says through clenched teeth, his voice low and raspy. "It's... just some visitors from the city-state of Nova Terra. They're here to trade something."

villager 38, gossip retelling villager 98 · pulled live, 2026-08-03

FathomPulse. honest glass. WIP

A dashboard on Fathom's live stream that measures its own truthfulness, not just the market's. Architecture and design system locked, no working pipeline yet.

STATUSwork in progress
LOCKED 2026-07-03
  • thesis: every panel carries a data-age chip, latency decomposed into provable stages, not smoothed away
  • frame semantics: one WS frame is a self-contained 10Hz snapshot, no client-side delta-stitching to get wrong
  • design tokens shipped and measured: waterfall ΔE 41.3, bid/ask ΔE 66.4, both pass colorblind-safe checks

pulse/ingest, pulse/api, pulse/web don't exist yet, this is genesis + design system only

THE FULL STORY ↓

Why this project. Fathom proves itself in a P99 number and a dashboard nobody but me has seen. Pulse is the part a recruiter can actually watch: Fathom’s internals turned into a screen, built on the same honesty standard as the pipeline underneath it. Most dashboards pretend to be real-time. Pulse would display its own staleness instead of hiding it.

Why this stack. It rides Fathom’s existing Kafka topics read-only, one more consumer off the same derived stream. FastAPI over WebSocket, one frame per 10Hz tick, each frame a complete snapshot of the trailing second so the client never has to stitch deltas together and get it wrong.

Honest status. Architecture and frame semantics locked, and the design system is more than a mood board: real contrast math run against the actual dark palette before a single component was built. Past that, there’s no ingest, no API, no web app. It stays labeled work-in-progress until it isn’t; this site doesn’t do vaporware.

WHAT I DIDN'T KNOW GOING IN

  • the honest-glass thesis came before a line of code did, on purpose
  • spent longer validating chart contrast math than writing a single line of ingest

ProvenanceFlow. where data comes from. WIP

FAIR-compliant data lineage tracker built on W3C PROV: tracing what touched your data, when, and why.

STATUSwork in progress
STANDARDW3C PROV
THE FULL STORY ↓

Why this project. I spent 18 months watching production data change shape between Bronze and Gold, and learned the hard way (see: the silent data loss incident) that lineage you can’t query is lineage you don’t have. FAIR data principles say provenance should be machine-actionable; almost nobody’s pipelines actually deliver that.

Why this stack. W3C PROV as the data model because standards outlive side projects. Lineage recorded in PROV today can become a Fair Digital Object tomorrow without rework.

Honest status. The PROV core works; the ambitions around it are still being built. It stays labeled work-in-progress until it isn’t; this site doesn’t do vaporware.

WHAT I DIDN'T KNOW GOING IN

  • read the W3C PROV spec twice, understood it the third time
  • learned the difference between a demo and a tool by building the wrong one first

EDUCATION

hover a degree, the coursework is a sky

M.Sc. Applied Data Science

Georg-August-Universität Göttingen · Göttingen, Germany

MAR 2026 — PRESENT first semester · in progress
Data Management Data Science Infrastructures Semistructured Data & XML German

B.Tech Computer Science & Engineering

Guru Gobind Singh Indraprastha University · Delhi, India

NOV 2021 — MAY 2025 GPA 9.62 / 10 (≈ 1.2 on the German scale, sehr gut) MINOR · ARTIFICIAL INTELLIGENCE & MACHINE LEARNING · +20 CREDITS
Discrete Mathematics Programming Fundamentals Python Programming Probability, Statistics & Linear Programming Data Structures Computational Methods Database Management Systems Object-Oriented Programming Automata Theory Design & Analysis of Algorithms Compiler Design Computer Networks Operating Systems Data Warehousing & Data Mining Data Modelling Statistics, Statistical Modelling & Data Analytics Artificial Intelligence Machine Learning Reinforcement Learning & Deep Learning Computer Vision Minor Project: CLIP vs YOLO vs CNN, allergen detection

COMPETENCIES

claims wired to receipts

Shipped with Indian colleagues for US clients; now studying in Germany with an international cohort. Promoted from intern to engineer in eleven months while learning on the run, and briefing leadership before turning 22. Independent by default, target-oriented by habit. The adjectives are yours to conclude; the facts are above.

stakeholder comms, up to C-suitecross-timezone teams · IN / US / DElearning on the runownership end-to-endexplaining tech to non-tech rooms
STREAMING PRODUCTION DE DATA QUALITY OBSERVABILITY MULTI-AGENT SIM LANGUAGES FATHOM FATHOMPULSE 10-BIT VILLAGE VEERSA
Streaming & real-time systems Kafka / Redpanda · Protobuf · TimescaleDB · ClickHouse → Fathom: P99 63.4 ms, live soak Production data engineering Databricks · Delta Lake · Spark · Snowflake · medallion · SCD 1/2 → 18 months in production at Veersa Data quality & reliability DQ validation · alerts · watermark ingestion → the silent data-loss incident: found, fixed, backfilled Observability Prometheus · Grafana · pipeline monitoring → every Fathom claim on this page is a dashboard first Multi-agent systems & simulation deterministic cores · checkpointing · agent design → 10-Bit Village: emergent and reproducible Languages Python · SQL · PySpark · C/C++ (DSA) → everything above GenAI-accelerated workflow Claude Code · LLM APIs · agent pipelines · prompt-driven tooling → this site, designed and built with an agent in the loop. human had final cut

THE SHELF

parked, not dead

parked 2025

Brain Decoder

Nishimoto 2011, but with modern text/video generation bolted on: reconstruct what the brain saw. POC code exists. Waiting for its month.

parked 2026

Haryanvi LLM

Fine-tune an open model on my mother tongue. Native speaker from Bahadurgarh; the dataset problem is the fun part.

incubating

Agent OS

Thesis: agent coordination is an operating-systems problem, and markdown files are the syscalls. Half the tooling already exists; I built it by accident.

RESEARCH

peer-reviewed

Kundra, R., & Ojaswi. “Assessing the Efficiency of Gradient Descent Variants in Training Neural Networks.” Darpan International Research Analysis, vol. 12, no. 3, pp. 596–604, Sep. 2024. DOI: 10.36676/dira.v12.i3.114

COLOPHON

the site, held to its own standard

A portfolio should meet the same bar as a pipeline: measured, honest, cheap to run. So here are this page’s own receipts.

TOTAL SHIPPED ~102 KB, all three pages smaller than one hero image on most portfolios
JAVASCRIPT ~11 KB the name animation and the dog included
FRAMEWORKS IN YOUR BROWSER 0 Astro compiles itself away
WEBFONT BYTES 0 system fonts only
COOKIES · TRACKERS · ANALYTICS 0 · 0 · 0 you were never the product here
CANVAS hand-rolled no animation libraries were harmed or used

✳ built in a dozen rounds of brutal feedback with an AI pair; the human had final cut. the dog is not a metaphor; he’s a labrador.