Skip to content
View andrebassi's full-sized avatar

Block or report andrebassi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
andrebassi/README.md

André Bassi

AI/LLM Engineer & Platform Engineer — RAG, agentes e LLMOps em produção · Kubernetes multi-cloud AWS·GCP·Azure · Go · Rust · Python

Construo a camada onde a IA roda em produção.

São 27 anos em infraestrutura — comecei em 1999 com ASP e PHP num servidor NT zumbindo do lado da mesa — e os últimos três dedicados a LLM em produção.

📍 Barueri, São Paulo · remoto · andrebassi.com.br · linkedin.com/in/andrebassi · contato@andrebassi.com.br


Atuação

apiVersion: engineering/v1
kind: PlatformEngineer
metadata:
  name: andre-bassi
  labels:
    focus: ai-infrastructure
    seniority: staff-principal
    availability: open-to-work
  annotations:
    site: https://andrebassi.com.br
    linkedin: https://linkedin.com/in/andrebassi
    email: contato@andrebassi.com.br

spec:
  role: Consultor · Platform & AI Engineering
  since: 2026-02                      # autônomo, andrebassi.com.br
  location: Barueri, São Paulo, Brasil
  remoto: true
  experiencia: 27 anos (desde 1999)

  competenciasPrincipais:             # o pin de 5 do LinkedIn
    - Large Language Models (LLM)
    - Kubernetes
    - Platform Engineer
    - Geração aumentada de recuperação (RAG)
    - SRE

  linguagens: [Go, Rust, Python, Shell]
  clouds:     [AWS, GCP, Azure, OCI]

  idiomas:
    portugues: nativo
    ingles: profissional completo

  abertoA:
    - Staff / Principal Platform Engineer
    - AI Infrastructure
    - Cloud Architect

status:
  phase: Running
  uptime: 27y                         # 1999 → hoje, sem gap
  conditions:
    - type: LLMOpsEmProducao
      status: "True"
      reason: vLLM · NVIDIA Triton · NIM · ONNX · MLflow · Kubeflow · Ray
    - type: RAGEmProducao
      status: "True"
      reason: busca híbrida · re-ranking · pgvector · LangChain/LangGraph · MCP
    - type: KubernetesMultiCloud
      status: "True"
      reason: EKS · GKE · AKS · OCI · Talos · Istio · Kong · KEDA · ArgoCD
    - type: ObservabilidadeFimAFim
      status: "True"
      reason: OpenTelemetry · Prometheus · Grafana
    - type: FinOps
      status: "True"
      reason: custo medido por unidade de trabalho, não por fatura
    - type: GPUWorkloads
      status: "True"
      reason: de GPU em nuvem sob demanda a Jetson Orin NX embarcado
    - type: Certificacoes
      status: "False"
      reason: CKA em estudo                    # honestidade > badge

MLOps e LLMOps — pipelines RAG end-to-end com busca híbrida, re-ranking e pgvector; agentes e multiagentes com LangChain/LangGraph; servidores MCP; avaliação, versionamento de prompt, tracing de inferência e guardrails. Model serving com vLLM, NVIDIA Triton e NIM, runtime ONNX, MLflow, Kubeflow e Ray — sobre workloads GPU, de GPU em nuvem sob demanda a Jetson Orin NX embarcado, com offload camada a camada entre CPU e GPU. Integração com OpenAI, Anthropic (Claude), Amazon Bedrock e Vertex AI.

Plataforma e operação — Kubernetes multi-cloud em produção (EKS, GKE, AKS, OCI, Talos), Docker, Helm, Istio como service mesh, Kong como API Gateway, KEDA, GitOps com ArgoCD, CI/CD em GitHub Actions e GitLab CI, IaC em Terraform e Pulumi. Observabilidade fim a fim — métricas, logs e tracing distribuído com OpenTelemetry, Prometheus e Grafana. Redes, balanceamento de carga, alta disponibilidade multi-região, RBAC, secrets management e políticas de segurança em cloud. FinOps: custo medido por unidade de trabalho, não por fatura. Governança, auditoria e LGPD em ambiente bancário e enterprise.


O que construí, com os números medidos

Projeto O que é Números Stack
edgeProxy · docs Proxy TCP/HTTP de edge. Geo-routing (MaxMind + Anycast), replicação SQLite via SWIM + QUIC, OpenTelemetry nativo, multi-PoP (GRU · IAD · SIN · FRA) Substitui nginx + haproxy + lua por um binário de ~3 MB de RSS, p99 sub-milissegundo no L4. 490+ testes Rust
Runner Codes · runner.codes Sandbox para código gerado por LLM em microVMs Firecracker. MCP-ready: o modelo escreve, o runner executa, o humano valida ~125 ms de cold-boot, 40+ runtimes prontos. Open source sob MPL-2.0 Go · Firecracker · KVM
infra-operator · docs Operator que declara recursos AWS (EKS, RDS, S3, IAM, VPC, Route53) como CRDs nativos Reconciliação contínua, drift detection e rollback automático — IaC dentro do control loop Go · Kubernetes
AudioFlow · audioflow.pro SaaS em produção que transcreve áudio do WhatsApp com IA, sem depender da API oficial da Meta R$ 0,0072 de IA por minuto processado, sustentando 94–97% de margem bruta. Postmortem público do billing que cobrou 12× a mais Go hexagonal · Temporal · Next.js · Supabase
Booster K1 · deep-dive Cérebro de voz de um humanoide de 22 graus de liberdade, embarcado RAG de eventos de 93,5% → 100% de recall num Jetson Orin NX de 8 GB Go hexagonal · Jetson Orin NX
AutoDJ · autodj.andrebassi.com.br Tocador que mixa como DJ e anda sozinho — o áudio toca no navegador, o Go só mantém a agenda e decide a próxima faixa BPM medido por ffmpeg/aubio, separação de voz por IA (Demucs) a ~US$ 0,001 por faixa. MIT Go · Next.js · Web Audio · Demucs

Stack

Platform & Orchestration Kubernetes · Istio · Kong · KEDA · ArgoCD · Helm · Cilium · Talos

AI / LLM Infrastructure RAG · pgvector · LangChain · LangGraph · MCP · vLLM · NVIDIA Triton · NIM · ONNX · MLflow · Kubeflow · Ray · Bedrock · Vertex AI

Infrastructure as Code Terraform · Pulumi · Ansible · Crossplane

Observability & SRE OpenTelemetry · Prometheus · Grafana · FinOps

Languages Go · Rust · Python · Shell

Clouds AWS · GCP · Azure · OCI


🇺🇸 English

AI/LLM Engineer & Platform Engineer — production RAG, agents and LLMOps · multi-cloud Kubernetes AWS·GCP·Azure · Go · Rust · Python

I build the layer where AI runs in production.

27 years in infrastructure — I started in 1999 with ASP and PHP on an NT server humming next to my desk — and the last three focused on LLMs in production.

MLOps and LLMOps — end-to-end RAG pipelines with hybrid search, re-ranking and pgvector; agents and multi-agent systems with LangChain/LangGraph; MCP servers; evaluation, prompt versioning, inference tracing and guardrails. Model serving with vLLM, NVIDIA Triton and NIM, ONNX runtime, MLflow, Kubeflow and Ray — on GPU workloads, from on-demand cloud GPU to an embedded Jetson Orin NX, with layer-by-layer offload between CPU and GPU.

Platform and operations — multi-cloud Kubernetes in production (EKS, GKE, AKS, OCI, Talos), Docker, Helm, Istio as service mesh, Kong as API Gateway, KEDA, GitOps with ArgoCD, CI/CD on GitHub Actions and GitLab CI, IaC in Terraform and Pulumi. End-to-end observability with OpenTelemetry, Prometheus and Grafana. Networking, load balancing, multi-region high availability, RBAC, secrets management and cloud security policies. FinOps: cost measured per unit of work, not per invoice.

Project What it is Measured Stack
edgeProxy Edge TCP/HTTP proxy. Geo-routing (MaxMind + Anycast), SQLite replication over SWIM + QUIC, native OpenTelemetry, multi-PoP Replaces nginx + haproxy + lua with a single ~3 MB RSS binary, sub-millisecond p99 at L4. 490+ tests Rust
Runner Codes Sandbox for LLM-generated code on Firecracker microVMs. MCP-ready ~125 ms cold-boot, 40+ ready runtimes. Open source under MPL-2.0 Go · Firecracker
infra-operator Go operator declaring AWS resources (EKS, RDS, S3, IAM, VPC, Route53) as native CRDs Continuous reconciliation, drift detection and automatic rollback Go · Kubernetes
AudioFlow · audioflow.pro Production SaaS transcribing WhatsApp audio with AI R$ 0.0072 of AI per minute processed, 94–97% gross margin. Public postmortem of the billing bug that overcharged 12× Go · Temporal · Next.js
Booster K1 · deep-dive Voice brain of a 22-degrees-of-freedom humanoid robot Event RAG from 93.5% → 100% recall on an 8 GB Jetson Orin NX Go · Jetson Orin NX
AutoDJ Autonomous DJ player — audio plays in the browser, Go only keeps the schedule BPM via ffmpeg/aubio, AI voice separation (Demucs) at ~US$ 0.001 per track. MIT Go · Next.js · Web Audio

Open to Staff/Principal Platform Engineer, AI Infrastructure and Cloud Architect roles — remote or hybrid in São Paulo.


Aberto a Staff/Principal Platform Engineer, AI Infrastructure e Cloud Architect — remoto ou híbrido em São Paulo.

LinkedIn Website Email Twitter

@andrebassi's activity is private