Data FinOps
Specification-driven assessment and evidence-led reporting.
DATA · ML · AI PLATFORM
I build data platforms and AI systems that teams can inspect, reproduce, and operate. More than twelve years across software, data engineering, machine learning, and cloud delivery.

SELECTED WORK
Specification-driven assessment and evidence-led reporting.
Brazilian education and procurement data contracts for reproducible releases.
Two Kaggle datasets, version 1 each. PNCP procurement history: 1,979 rows per layer, notices from 2025-01-01 to 2025-01-07, modality 6. Education data lake: SIOPE annual municipal declarations 2019-2023 joined to IBGE codes, 27,830 municipality-year rows; a clean download was hash-verified.
Traceable municipality-level anomaly triage from public data.
v0.2.0 trains a robust peer-group anomaly-triage model on the pinned education Kaggle release. The 2022 batch was scored (96 review signals); the drift gate blocked the 2023 batch (spread ratio 1.384).
Outputs are review signals only. There are no labels, so no accuracy is claimed.
Transparent retrieval and ranking with responsible matching limits.
v0.2.0 ranks historical notices from the pinned PNCP Kaggle release and verifies its hash before use.
Offline evaluation uses synthetic profiles with rule-derived judgments, so it measures constraint adherence, not user relevance.
Evaluated retrieval, source citations, and redacted traces.
Redis-coordinated workers, Kubernetes, Terraform, and recovery tests.
Local Compose benchmark on 2026-09-24: 48 requests at concurrency 6, 61.89 req/s, p95 147.95 ms. Worker recovery proven on a local kind cluster.
Deterministic model stub on one Docker Desktop host; not a capacity claim. Cloud not validated.
CAPABILITIES
Data contracts, reproducible public releases, and cost assessment.
Traceable pipelines, offline evaluation, and transparent ranking.
Retrieval evaluation, source citations, and redacted traces.
Redis coordination, Kubernetes, Terraform, and recovery tests.
EXPERIENCE
2024 — present
Data pipelines, AI applications, governance, and observability.
2023 — 2024
Lakehouse pipelines, MLflow monitoring, and versioned data assets.
2022 — 2023
Composer, Dataproc, Vertex AI, and CI/CD for data systems.
2014 — 2022
Data platforms, full-stack systems, databases, and embedded software.
WORKING PRINCIPLES
I use reproducible data releases, explicit evaluation, infrastructure-as-code, and documented limits to make technical work reviewable.
CONTACT