AI/LLM Engineer & Platform Engineer — RAG, agentes e LLMOps em produção · Kubernetes multi-cloud AWS·GCP·Azure · Go · Rust · Python
Construo a camada onde a IA roda em produção.
São 27 anos em infraestrutura — comecei em 1999 com ASP e PHP num servidor NT zumbindo do lado da mesa — e os últimos três dedicados a LLM em produção.
📍 Barueri, São Paulo · remoto · andrebassi.com.br · linkedin.com/in/andrebassi · contato@andrebassi.com.br
apiVersion: engineering/v1
kind: PlatformEngineer
metadata:
name: andre-bassi
labels:
focus: ai-infrastructure
seniority: staff-principal
availability: open-to-work
annotations:
site: https://andrebassi.com.br
linkedin: https://linkedin.com/in/andrebassi
email: contato@andrebassi.com.br
spec:
role: Consultor · Platform & AI Engineering
since: 2026-02 # autônomo, andrebassi.com.br
location: Barueri, São Paulo, Brasil
remoto: true
experiencia: 27 anos (desde 1999)
competenciasPrincipais: # o pin de 5 do LinkedIn
- Large Language Models (LLM)
- Kubernetes
- Platform Engineer
- Geração aumentada de recuperação (RAG)
- SRE
linguagens: [Go, Rust, Python, Shell]
clouds: [AWS, GCP, Azure, OCI]
idiomas:
portugues: nativo
ingles: profissional completo
abertoA:
- Staff / Principal Platform Engineer
- AI Infrastructure
- Cloud Architect
status:
phase: Running
uptime: 27y # 1999 → hoje, sem gap
conditions:
- type: LLMOpsEmProducao
status: "True"
reason: vLLM · NVIDIA Triton · NIM · ONNX · MLflow · Kubeflow · Ray
- type: RAGEmProducao
status: "True"
reason: busca híbrida · re-ranking · pgvector · LangChain/LangGraph · MCP
- type: KubernetesMultiCloud
status: "True"
reason: EKS · GKE · AKS · OCI · Talos · Istio · Kong · KEDA · ArgoCD
- type: ObservabilidadeFimAFim
status: "True"
reason: OpenTelemetry · Prometheus · Grafana
- type: FinOps
status: "True"
reason: custo medido por unidade de trabalho, não por fatura
- type: GPUWorkloads
status: "True"
reason: de GPU em nuvem sob demanda a Jetson Orin NX embarcado
- type: Certificacoes
status: "False"
reason: CKA em estudo # honestidade > badgeMLOps e LLMOps — pipelines RAG end-to-end com busca híbrida, re-ranking e pgvector; agentes e multiagentes com LangChain/LangGraph; servidores MCP; avaliação, versionamento de prompt, tracing de inferência e guardrails. Model serving com vLLM, NVIDIA Triton e NIM, runtime ONNX, MLflow, Kubeflow e Ray — sobre workloads GPU, de GPU em nuvem sob demanda a Jetson Orin NX embarcado, com offload camada a camada entre CPU e GPU. Integração com OpenAI, Anthropic (Claude), Amazon Bedrock e Vertex AI.
Plataforma e operação — Kubernetes multi-cloud em produção (EKS, GKE, AKS, OCI, Talos), Docker, Helm, Istio como service mesh, Kong como API Gateway, KEDA, GitOps com ArgoCD, CI/CD em GitHub Actions e GitLab CI, IaC em Terraform e Pulumi. Observabilidade fim a fim — métricas, logs e tracing distribuído com OpenTelemetry, Prometheus e Grafana. Redes, balanceamento de carga, alta disponibilidade multi-região, RBAC, secrets management e políticas de segurança em cloud. FinOps: custo medido por unidade de trabalho, não por fatura. Governança, auditoria e LGPD em ambiente bancário e enterprise.
| Projeto | O que é | Números | Stack |
|---|---|---|---|
| edgeProxy · docs | Proxy TCP/HTTP de edge. Geo-routing (MaxMind + Anycast), replicação SQLite via SWIM + QUIC, OpenTelemetry nativo, multi-PoP (GRU · IAD · SIN · FRA) | Substitui nginx + haproxy + lua por um binário de ~3 MB de RSS, p99 sub-milissegundo no L4. 490+ testes |
Rust |
| Runner Codes · runner.codes | Sandbox para código gerado por LLM em microVMs Firecracker. MCP-ready: o modelo escreve, o runner executa, o humano valida | ~125 ms de cold-boot, 40+ runtimes prontos. Open source sob MPL-2.0 | Go · Firecracker · KVM |
| infra-operator · docs | Operator que declara recursos AWS (EKS, RDS, S3, IAM, VPC, Route53) como CRDs nativos | Reconciliação contínua, drift detection e rollback automático — IaC dentro do control loop | Go · Kubernetes |
| AudioFlow · audioflow.pro | SaaS em produção que transcreve áudio do WhatsApp com IA, sem depender da API oficial da Meta | R$ 0,0072 de IA por minuto processado, sustentando 94–97% de margem bruta. Postmortem público do billing que cobrou 12× a mais | Go hexagonal · Temporal · Next.js · Supabase |
| Booster K1 · deep-dive | Cérebro de voz de um humanoide de 22 graus de liberdade, embarcado | RAG de eventos de 93,5% → 100% de recall num Jetson Orin NX de 8 GB | Go hexagonal · Jetson Orin NX |
| AutoDJ · autodj.andrebassi.com.br | Tocador que mixa como DJ e anda sozinho — o áudio toca no navegador, o Go só mantém a agenda e decide a próxima faixa | BPM medido por ffmpeg/aubio, separação de voz por IA (Demucs) a ~US$ 0,001 por faixa. MIT |
Go · Next.js · Web Audio · Demucs |
Platform & Orchestration Kubernetes · Istio · Kong · KEDA · ArgoCD · Helm · Cilium · Talos
AI / LLM Infrastructure RAG · pgvector · LangChain · LangGraph · MCP · vLLM · NVIDIA Triton · NIM · ONNX · MLflow · Kubeflow · Ray · Bedrock · Vertex AI
Infrastructure as Code Terraform · Pulumi · Ansible · Crossplane
Observability & SRE OpenTelemetry · Prometheus · Grafana · FinOps
Languages Go · Rust · Python · Shell
Clouds AWS · GCP · Azure · OCI
🇺🇸 English
AI/LLM Engineer & Platform Engineer — production RAG, agents and LLMOps · multi-cloud Kubernetes AWS·GCP·Azure · Go · Rust · Python
I build the layer where AI runs in production.
27 years in infrastructure — I started in 1999 with ASP and PHP on an NT server humming next to my desk — and the last three focused on LLMs in production.
MLOps and LLMOps — end-to-end RAG pipelines with hybrid search, re-ranking and pgvector; agents and multi-agent systems with LangChain/LangGraph; MCP servers; evaluation, prompt versioning, inference tracing and guardrails. Model serving with vLLM, NVIDIA Triton and NIM, ONNX runtime, MLflow, Kubeflow and Ray — on GPU workloads, from on-demand cloud GPU to an embedded Jetson Orin NX, with layer-by-layer offload between CPU and GPU.
Platform and operations — multi-cloud Kubernetes in production (EKS, GKE, AKS, OCI, Talos), Docker, Helm, Istio as service mesh, Kong as API Gateway, KEDA, GitOps with ArgoCD, CI/CD on GitHub Actions and GitLab CI, IaC in Terraform and Pulumi. End-to-end observability with OpenTelemetry, Prometheus and Grafana. Networking, load balancing, multi-region high availability, RBAC, secrets management and cloud security policies. FinOps: cost measured per unit of work, not per invoice.
| Project | What it is | Measured | Stack |
|---|---|---|---|
| edgeProxy | Edge TCP/HTTP proxy. Geo-routing (MaxMind + Anycast), SQLite replication over SWIM + QUIC, native OpenTelemetry, multi-PoP | Replaces nginx + haproxy + lua with a single ~3 MB RSS binary, sub-millisecond p99 at L4. 490+ tests |
Rust |
| Runner Codes | Sandbox for LLM-generated code on Firecracker microVMs. MCP-ready | ~125 ms cold-boot, 40+ ready runtimes. Open source under MPL-2.0 | Go · Firecracker |
| infra-operator | Go operator declaring AWS resources (EKS, RDS, S3, IAM, VPC, Route53) as native CRDs | Continuous reconciliation, drift detection and automatic rollback | Go · Kubernetes |
| AudioFlow · audioflow.pro | Production SaaS transcribing WhatsApp audio with AI | R$ 0.0072 of AI per minute processed, 94–97% gross margin. Public postmortem of the billing bug that overcharged 12× | Go · Temporal · Next.js |
| Booster K1 · deep-dive | Voice brain of a 22-degrees-of-freedom humanoid robot | Event RAG from 93.5% → 100% recall on an 8 GB Jetson Orin NX | Go · Jetson Orin NX |
| AutoDJ | Autonomous DJ player — audio plays in the browser, Go only keeps the schedule | BPM via ffmpeg/aubio, AI voice separation (Demucs) at ~US$ 0.001 per track. MIT |
Go · Next.js · Web Audio |
Open to Staff/Principal Platform Engineer, AI Infrastructure and Cloud Architect roles — remote or hybrid in São Paulo.
Aberto a Staff/Principal Platform Engineer, AI Infrastructure e Cloud Architect — remoto ou híbrido em São Paulo.