LM

Luis Mori Guerra

Engineering Manager / Technical Lead

HomeAbout MeTopics

Recent Articles

Anticipate the Path: The Simple Reason Rules, Skills, and Agents WorkHow Long Should an LLM Response Be? Deriving a Length Budget From Human Reading CapacityModern Algorithms Every Programmer Should Know: From Matching Markets to HNSW and ARCOh My Pi (omp.sh): A Quick Look at the Hardened Coding Agent CLIHow to Build the Fastest, Lightest Autonomous Agent Platform with Kubernetes, ArgoCD, Ante, and Frontier LLMs

Topics

AI Agents77AI30Architecture27LLM23Claude Code22Cursor21AI Engineering20Loop Engineering20
luismori.
ArticlesAboutTopicsGitHubConnect
← All topics

Inference

A topic hub collecting every article tagged Inference. Use it to explore related posts and follow this theme across the site.

2 articles

Explore More Topics

AI Agents AI Architecture LLM Claude Code Cursor
AI LLM AI Engineering OpenAI Inference Benchmarks Open Source

How Fast Is gpt-oss-120b? Speed, Quality, and Routing Tradeoffs in 2026

For engineers evaluating gpt-oss-120b: OpenAI's open-weight model is fast for an open reasoning checkpoint, but throughput depends heavily on the provider and benchmark view. Here is how to read Artificial Analysis, OpenRouter trends, and the cost tradeoffs behind fast inference.

Jul 3, 2026 20 min read
AI Infrastructure MLOps Kubernetes vLLM KServe GPU Inference Engineering

ML/AI Infrastructure: A Quick Look, and the Skills You Need

A hands-on, source-backed tour of the modern ML/AI infrastructure stack, centered on self-hosting LLM inference on Kubernetes with vLLM and KServe, plus the GPU economics and skills that actually matter in 2026.

Jun 16, 2026 17 min read

Quick find

Search the blog

Search by topic, title, framework, or pattern.

Interactive Diagram Viewer 100%
Tip: Drag to pan • Scroll to zoom • Double-click to reset • Press Esc to exit