Data & AI Engineer — batch and streaming pipelines (BigQuery, dbt, Airflow, PySpark) and production LLM applications, including work on the Claude API.
Currently: building out end-to-end data platforms with governance and IaC baked in, not bolted on.
Stack
| Ingestion & Processing | Warehouse & Modeling | Orchestration & Infra | AI / LLM |
|---|---|---|---|
| Python, PySpark, pandas | BigQuery, dbt, SQL | Airflow, Docker, Terraform | Claude API, RAG, MCP |
Featured projects
- github-firehose-activity — PySpark pipeline processing 100M+ GitHub Archive events into a BigQuery star schema, dbt-tested, Airflow-orchestrated
- terraform-gcp-data-platform — A GCP analytics platform where every resource is IaC, and an AI viewer on top is infrastructurally incapable of writing data
- credit-risk-explainable-ai — Credit-risk scoring (XGBoost + SHAP) with Claude-generated, FCRA-style explanations and fair-lending guardrails
- retrieval-benchmark — Measured retrieval quality: naive vector search vs. hybrid BM25/vector + rerank, over a real 10-K filing
- b2-mcp-server — MCP server exposing Backblaze B2 as tools an AI client can call, with a fenced filesystem and audit-logged deletes
- b2-sync — Backup CLI in Go with SHA-256 diffing and a bounded worker pool
📫 LinkedIn · felipefumero.com



