✕ Clear all filters
546 articles
▶ Videos →

MLOps & LLMOps Reads

546 articles · Updated every 3 hours · View all reads

All Articles 189,225Blog Posts 171,155Tech Tutorials 50,720Research Papers 36,770News 23,155 ⚡ AI Lessons
MLOps and LLMOps: How AI Systems Are Managed in Practice
Medium · Data Science 🏭 MLOps & LLMOps 1w ago
MLOps and LLMOps: How AI Systems Are Managed in Practice
Creating an artificial intelligence model is only one part of building a useful AI system. After development, the model must be tested… Continue reading on Medi
LLM Inference — 4/7 [The Software Stack Behind High-Performance LLM Inference]
Medium · LLM 🏭 MLOps & LLMOps 1w ago
LLM Inference — 4/7 [The Software Stack Behind High-Performance LLM Inference]
From CUDA kernels and PyTorch to vLLM, SGLang, TensorRT-LLM, distributed serving, and performance benchmarking Continue reading on Medium »
Dev.to AI 🏭 MLOps & LLMOps ⚡ AI Lesson 1w ago
Power Platform Pipelines - Moving Flows from Dev to Prod with Approvals
"How do you deploy flows to production?" I ask this question in every governance workshop. The most common answer is still "export as managed, import manually,
Why AI Models Break in Production — Data Drift, Model Drift & Concept Drift Explained
Medium · Data Science 🏭 MLOps & LLMOps 1w ago
Why AI Models Break in Production — Data Drift, Model Drift & Concept Drift Explained
Your team spent three months building a fraud detection model. Continue reading on Medium »
Dev.to AI 🏭 MLOps & LLMOps ⚡ AI Lesson 1w ago
How many tools should an MCP server have?
If you maintain an MCP server, at some point you ask this. You've got twenty tools, you're about to add five more, and something feels wrong about it — but you
How to Use unsloth/Qwen3.8–27B-GGUF in Claude Code via Ollama Without Dying in the Process? (2/2)
Medium · LLM 🏭 MLOps & LLMOps 3w ago
How to Use unsloth/Qwen3.8–27B-GGUF in Claude Code via Ollama Without Dying in the Process? (2/2)
Fitting 27 billion parameters, a context window that’s actually useful, and an overthinking agent into 24 GB of VRAM Continue reading on Towards AI »
Your Model Is Ready. Your Release Path Is Still a Ritual
Medium · DevOps 🏭 MLOps & LLMOps ⚡ AI Lesson 3w ago
Your Model Is Ready. Your Release Path Is Still a Ritual
Originally published on TruFyre. Continue reading on Medium »
TinyML Deployment: Quantization, Pruning & Memory Optimization for Microcontrollers
Dev.to · beefed.ai 🏭 MLOps & LLMOps 3w ago
TinyML Deployment: Quantization, Pruning & Memory Optimization for Microcontrollers
Practical guide to quantization, pruning, and memory techniques to run accurate ML models on microcontrollers with TinyML.
Stop Using Ollama in Production: Why vLLM & SGLang Win
Medium · Machine Learning 🏭 MLOps & LLMOps 3w ago
Stop Using Ollama in Production: Why vLLM & SGLang Win
The infrastructure you use to prototype is quietly sabotaging your production agents. Here’s how to fix the architectural mismatch. Continue reading on Medium »
LLM Inference Servers simply explained
Medium · Machine Learning 🏭 MLOps & LLMOps 4w ago
LLM Inference Servers simply explained
Through the example of vLLM and TensorRT-LLM. Continue reading on Medium »
LLM Inference Servers simply explained
Medium · LLM 🏭 MLOps & LLMOps 4w ago
LLM Inference Servers simply explained
Through the example of vLLM and TensorRT-LLM. Continue reading on Medium »
Two Acquisitions Deep, Two Still Standing: What's Really Happening to Open-Source LLM Observability
Hackernoon 🏭 MLOps & LLMOps 1mo ago
Two Acquisitions Deep, Two Still Standing: What's Really Happening to Open-Source LLM Observability
Langfuse and Helicone were acquired in 2026 while Opik and Phoenix stayed independent — here's what each acquirer actually promised in writing.
Desmistificando o Deploy de LLMs em 2026: Quando usar Ollama ou vLLM (e por que não são a mesma…
Medium · LLM 🏭 MLOps & LLMOps 1mo ago
Desmistificando o Deploy de LLMs em 2026: Quando usar Ollama ou vLLM (e por que não são a mesma…
Olá, pessoal. Continue reading on Medium »
¿Cómo usar unsloth/Qwen3.8–27B-GGUF
Medium · LLM 🏭 MLOps & LLMOps 1mo ago
¿Cómo usar unsloth/Qwen3.8–27B-GGUF
Son las once de la noche. Tienes 27 mil millones de parámetros descansando cómodamente en tu GPU, Claude Code instalado y ollama serve… Continue reading on Medi
Medium · Machine Learning 🏭 MLOps & LLMOps ⚡ AI Lesson 1mo ago
MLOPS LIFE CYCLE
MLOps, short for machine learning operations, is a set of practice employed to make machine learning model development deployable, usable… Continue reading on M
Medium · DevOps 🏭 MLOps & LLMOps ⚡ AI Lesson 1mo ago
MLOPS LIFE CYCLE
MLOps, short for machine learning operations, is a set of practice employed to make machine learning model development deployable, usable… Continue reading on M
¿Cómo usar unsloth/Qwen3.8–27B-GGUF
Medium · LLM 🏭 MLOps & LLMOps 1mo ago
¿Cómo usar unsloth/Qwen3.8–27B-GGUF
Son las once de la noche. Tienes 27 mil millones de parámetros descansando cómodamente en tu GPU, Claude Code instalado y ollama serve… Continue reading on Lati
DORA Metrics for ML Model Deployment: How Software Delivery Performance Applies to MLOps Pipelines
Medium · DevOps 🏭 MLOps & LLMOps 1mo ago
DORA Metrics for ML Model Deployment: How Software Delivery Performance Applies to MLOps Pipelines
DORA Metrics for ML Model Deployment: How Software Delivery Performance Applies to MLOps Pipelines Continue reading on Women in Technology »
Qwen3.8–27B-FP8 on Your Mac: 4 Ways to Run It (and When You Shouldn’t)
Medium · LLM 🏭 MLOps & LLMOps 1mo ago
Qwen3.8–27B-FP8 on Your Mac: 4 Ways to Run It (and When You Shouldn’t)
Complete guide to local deployment — MLX, GGUF, Ollama, LM Studio — plus the cloud alternative that might save you hours. Continue reading on Medium »
Medium · Machine Learning 🏭 MLOps & LLMOps ⚡ AI Lesson 1mo ago
Production'da ML Modeli Öldüğünde Kim Fark Eder? — Drift Tespiti ve Otomatik Retrain Pipeline'ı
PSI + KS testleri, FastAPI, Prometheus, MLflow ve auto-retrain ile uçtan uca mini MLOps mimarisi Continue reading on Medium »