Data engineer. I build recurring collection pipelines against public sources,
and the thing I care most about is the failure mode where a pipeline keeps
reporting success while returning nothing.
Recent: a daily multi-operator price-tracking pipeline with immutable dated
snapshots and coverage guards that fail the run rather than write an empty
file; an ESP32 WiFi CSI array doing RF fingerprinting from clock-offset
features (99.7% device ID over 7,497 windows on one receiver, 95.7% over
14,234 once a second receiver joins — a cross-receiver generalization gap);
scrapers against county permit portals.
Operations background before this, which is why I am comfortable being the
person in the room who can talk to the engineers and to the people paying for
the thing.
I build data pipelines that know when they're lying to you.
GAScout (https://gascout.pages.dev) — Georgia public-records pipeline. Four counties live, each a different ingestion problem resolved into one schema: a 9,915-page fixed-width mainframe PDF (409,142 records, zero dropped rows), a weekly XLSX roll, monthly text PDFs, and scanned sheriff's levy documents via OCR. Adding a county is a config file, not code. The other 155 are labeled planned rather than filled with invented numbers.
FindStorage (https://findstorage.pages.dev) — national self-storage pricing tracker. 4,639 locations, 245,000+ price changes logged, running unattended on a daily schedule. Aborts rather than publishing when the store count moves more than it should, because the failure that costs you isn't a crash — it's a scraper quietly returning less after a site changes and nobody noticing for three weeks.
Also built a WiFi CSI device-fingerprinting array on ESP32 — custom ESP-IDF firmware and a Python DSP pipeline, running as my apartment's security system. Cross-manufacturer separation holds at 11-15 sigma; same-model discrimination tops out near 77% and true clock twins are the open problem.
Before this I ran a 500-acre motocross park in Northern California: acquired dormant land, re-entitled it, directed a 120-person race-day crew, ran the medical plan through a mid-event helicopter evacuation, and sold my stake.
Self-taught, ~18 months in. Interested in data engineering, pipeline work, and anything where messy real-world sources have to become trustworthy structured data.
Remote: Yes
Willing to relocate: Yes
Technologies: Python, SQL, SQLite, Parquet, pandas, GitHub Actions, ESP-IDF/C, DSP (RANSAC phase-slope fitting), web scraping, ETL
Résumé/CV: https://braedenkeena.pages.dev
Email: Braeden@thekeenas.com
Data engineer. I build recurring collection pipelines against public sources, and the thing I care most about is the failure mode where a pipeline keeps reporting success while returning nothing.
Recent: a daily multi-operator price-tracking pipeline with immutable dated snapshots and coverage guards that fail the run rather than write an empty file; an ESP32 WiFi CSI array doing RF fingerprinting from clock-offset features (99.7% device ID over 7,497 windows on one receiver, 95.7% over 14,234 once a second receiver joins — a cross-receiver generalization gap); scrapers against county permit portals.
Operations background before this, which is why I am comfortable being the person in the room who can talk to the engineers and to the people paying for the thing.
reply