Construo a camada onde a
IA roda em produção.
I build the layer where
AI runs in production.
São 27 anos em infraestrutura — comecei em 1999 com ASP e PHP num servidor NT zumbindo do lado da mesa — e os últimos três dedicados a LLM em produção. Model serving, agentes, RAG e a plataforma que sustenta os três.
27 years in infrastructure — I started in 1999 with ASP and PHP on an NT box humming beside the desk — and the last three on LLM in production. Model serving, agents, RAG, and the platform underneath all three.
1999 ▸ 2026
1999 ▸ 2026
Seis eras de infraestrutura, sem pular nenhuma.
Six eras of infrastructure, none of them skipped.
Servidor físico
Bare metal
Sysadmin, redes e os primeiros datacenters em Linux e NT. ASP, PHP e VBScript — anos antes de existir nuvem para alugar.
Sysadmin, networks, and the first Linux/NT datacenters. ASP, PHP and VBScript — years before there was any cloud to rent.
Virtualização e a nuvem como promessa
Virtualization, and the cloud as a promise
.NET, e AWS, GCP e Azure quando ainda era aposta — não plataforma.
.NET, plus AWS, GCP and Azure back when they were a bet, not a platform.
Containers, depois Kubernetes
Containers, then Kubernetes
Docker e, na sequência, Kubernetes em produção — EKS, GKE, AKS, OCI e Talos. Istio, KEDA, GitOps com ArgoCD.
Docker, then Kubernetes in production — EKS, GKE, AKS, OCI and Talos. Istio, KEDA, GitOps with ArgoCD.
Transformer, BERT, RAG
Transformer, BERT, RAG
IA deixando de ser pesquisa e virando engenharia — com os mesmos problemas de sempre: latência, custo e o que fazer quando cai.
AI turning from research into engineering — with the same old problems: latency, cost, and what to do when it falls over.
LLM em produção
LLM in production
Model serving com vLLM, Triton e NIM sobre GPU — da nuvem sob demanda ao Jetson Orin NX embarcado. Agentes com LangChain e LangGraph, servidores MCP, avaliação e guardrails.
Model serving with vLLM, Triton and NIM on GPU — from on-demand cloud to an embedded Jetson Orin NX. Agents with LangChain and LangGraph, MCP servers, evals and guardrails.
Consultoria em AI Platform
AI Platform consulting
LLMOps, model serving e FinOps: custo medido por unidade de trabalho, não por fatura.
LLMOps, model serving and FinOps: cost measured per unit of work, not per invoice.
o que construí
what I built
Cada um com o número que ele mede.
Each one with the number it measures.
edgeProxy
Proxy TCP/HTTP de edge em Rust. Substitui nginx + haproxy + lua por um binário só. Geo-routing, replicação SQLite via SWIM + QUIC, multi-PoP. 490+ testes.
Edge TCP/HTTP proxy in Rust. Replaces nginx + haproxy + lua with a single binary. Geo-routing, SQLite replication over SWIM + QUIC, multi-PoP. 490+ tests.
RustQUICMaxMindOpenTelemetry ~125 ms de cold-boot · 40+ runtimescold boot · 40+ runtimesRunner Codes
Sandbox para código gerado por LLM em microVMs Firecracker. MCP-ready, open source sob MPL-2.0.
Sandbox for LLM-generated code on Firecracker microVMs. MCP-ready, open source under MPL-2.0.
FirecrackerGoMCPKVM 93,5% → 100% de recall no RAG de eventosevent-RAG recallBooster K1
Cérebro de voz de um humanoide de 22 graus de liberdade, num Jetson Orin NX de 8 GB. Go hexagonal, embarcado.
The voice brain of a 22-DoF humanoid, on an 8 GB Jetson Orin NX. Hexagonal Go, embedded.
GoJetson Orin NXROS2RAG R$ 0,0072 por minuto transcritoper minute transcribedAudioFlow
SaaS em produção que transcreve áudio do WhatsApp. Inclui o postmortem público do billing que cobrou 12× a mais por três meses.
A production SaaS that transcribes WhatsApp audio. Includes the public postmortem of the billing bug that overcharged 12× for three months.
Go hexagonalTemporalNext.jsSupabase 6 recursos AWS6 AWS resources como CRD nativoas native CRDsinfra-operator
Operator em Go que declara EKS, RDS, S3, IAM, VPC e Route53 como recursos do cluster — com reconciliação contínua, drift detection e rollback.
A Go operator that declares EKS, RDS, S3, IAM, VPC and Route53 as cluster resources — with continuous reconciliation, drift detection and rollback.
Gocontroller-runtimeAWSCRDa stack, por camada
the stack, by layer
Do modelo ao cabo.
From the model down to the cable.
Consultoria em AI Platform, LLMOps, model serving e Kubernetes — remoto ou híbrido em São Paulo.
Consulting on AI Platform, LLMOps, model serving and Kubernetes — remote or hybrid in São Paulo.
falar comigo → get in touch →