Construo a camada onde a
IA roda em produção.

I build the layer where
AI runs in production.

São 27 anos em infraestrutura — comecei em 1999 com ASP e PHP num servidor NT zumbindo do lado da mesa — e os últimos três dedicados a LLM em produção. Model serving, agentes, RAG e a plataforma que sustenta os três.

27 years in infrastructure — I started in 1999 with ASP and PHP on an NT box humming beside the desk — and the last three on LLM in production. Model serving, agents, RAG, and the platform underneath all three.

0anos em infraestruturayears in infrastructure
0anos de LLM em produçãoyears of LLM in production
0projetos com número medidoprojects with a measured number
0testes no edgeProxytests in edgeProxy

1999 ▸ 2026

1999 ▸ 2026

Seis eras de infraestrutura, sem pular nenhuma.

Six eras of infrastructure, none of them skipped.

Servidor físico

Bare metal

Sysadmin, redes e os primeiros datacenters em Linux e NT. ASP, PHP e VBScript — anos antes de existir nuvem para alugar.

Sysadmin, networks, and the first Linux/NT datacenters. ASP, PHP and VBScript — years before there was any cloud to rent.

Virtualização e a nuvem como promessa

Virtualization, and the cloud as a promise

.NET, e AWS, GCP e Azure quando ainda era aposta — não plataforma.

.NET, plus AWS, GCP and Azure back when they were a bet, not a platform.

Containers, depois Kubernetes

Containers, then Kubernetes

Docker e, na sequência, Kubernetes em produção — EKS, GKE, AKS, OCI e Talos. Istio, KEDA, GitOps com ArgoCD.

Docker, then Kubernetes in production — EKS, GKE, AKS, OCI and Talos. Istio, KEDA, GitOps with ArgoCD.

Transformer, BERT, RAG

Transformer, BERT, RAG

IA deixando de ser pesquisa e virando engenharia — com os mesmos problemas de sempre: latência, custo e o que fazer quando cai.

AI turning from research into engineering — with the same old problems: latency, cost, and what to do when it falls over.

LLM em produção

LLM in production

Model serving com vLLM, Triton e NIM sobre GPU — da nuvem sob demanda ao Jetson Orin NX embarcado. Agentes com LangChain e LangGraph, servidores MCP, avaliação e guardrails.

Model serving with vLLM, Triton and NIM on GPU — from on-demand cloud to an embedded Jetson Orin NX. Agents with LangChain and LangGraph, MCP servers, evals and guardrails.

Consultoria em AI Platform

AI Platform consulting

LLMOps, model serving e FinOps: custo medido por unidade de trabalho, não por fatura.

LLMOps, model serving and FinOps: cost measured per unit of work, not per invoice.

o que construí

what I built

Cada um com o número que ele mede.

Each one with the number it measures.

~3 MB de RSS · p99 sub-ms no L4of RSS · sub-ms p99 at L4

edgeProxy

Proxy TCP/HTTP de edge em Rust. Substitui nginx + haproxy + lua por um binário só. Geo-routing, replicação SQLite via SWIM + QUIC, multi-PoP. 490+ testes.

Edge TCP/HTTP proxy in Rust. Replaces nginx + haproxy + lua with a single binary. Geo-routing, SQLite replication over SWIM + QUIC, multi-PoP. 490+ tests.

RustQUICMaxMindOpenTelemetry
~125 ms de cold-boot · 40+ runtimescold boot · 40+ runtimes

Runner Codes

Sandbox para código gerado por LLM em microVMs Firecracker. MCP-ready, open source sob MPL-2.0.

Sandbox for LLM-generated code on Firecracker microVMs. MCP-ready, open source under MPL-2.0.

FirecrackerGoMCPKVM
93,5% → 100% de recall no RAG de eventosevent-RAG recall

Booster K1

Cérebro de voz de um humanoide de 22 graus de liberdade, num Jetson Orin NX de 8 GB. Go hexagonal, embarcado.

The voice brain of a 22-DoF humanoid, on an 8 GB Jetson Orin NX. Hexagonal Go, embedded.

GoJetson Orin NXROS2RAG
R$ 0,0072 por minuto transcritoper minute transcribed

AudioFlow

SaaS em produção que transcreve áudio do WhatsApp. Inclui o postmortem público do billing que cobrou 12× a mais por três meses.

A production SaaS that transcribes WhatsApp audio. Includes the public postmortem of the billing bug that overcharged 12× for three months.

Go hexagonalTemporalNext.jsSupabase
6 recursos AWS6 AWS resources como CRD nativoas native CRDs

infra-operator

Operator em Go que declara EKS, RDS, S3, IAM, VPC e Route53 como recursos do cluster — com reconciliação contínua, drift detection e rollback.

A Go operator that declares EKS, RDS, S3, IAM, VPC and Route53 as cluster resources — with continuous reconciliation, drift detection and rollback.

Gocontroller-runtimeAWSCRD

a stack, por camada

the stack, by layer

Do modelo ao cabo.

From the model down to the cable.

IA / LLMAI / LLM RAG híbridopgvectorre-rankingLangChainLangGraphMCPguardrailsLoRA · QLoRA
Model servingModel serving vLLMNVIDIA TritonNIMONNXMLflowKubeflowRayJetson Orin NX
PlataformaPlatform EKSGKEAKSOCITalosIstioKEDAArgoCDHelm
Infra como códigoInfra as code TerraformPulumiGitHub ActionsGitLab CIAnsible
Observar e pagarObserve and pay OpenTelemetryPrometheusGrafanaFinOps por unidade
LinguagensLanguages GoRustPythonBashTypeScript

Consultoria em AI Platform, LLMOps, model serving e Kubernetes — remoto ou híbrido em São Paulo.

Consulting on AI Platform, LLMOps, model serving and Kubernetes — remote or hybrid in São Paulo.

falar comigo → get in touch →