YC Yc.W home Systems · C++ · AI Infrastructure

LLM Articles

9 articles


AI & LLM · curated-v1

LLM 推理服务上线清单:从单机验证到多节点运维

Curated guide
整合 vLLM、Ray、Docker、模型网关和离线部署笔记,给出模型服务从容量估算到灰度、监控与回滚的生产检查表。
AI & LLM · curated-v1

AI Agent 工程化学习地图:从模型调用到可控协作

Curated guide
把 RAG、LangChain、LangGraph、DeepAgents、OpenCode 和模型网关等历史文章整理成一条从 API 调用到可观测 Agent 系统的学习路线。
AI & LLM

从零到一:构建你的 LLM 应用开发知识体系

Learning roadmap for LLM application development, mapping RAG, LangChain, LangGraph, and DeepAgents from fundamentals to production-grade agent workflows.
AI & LLM

从零搭建 vLLM 多节点推理集群:Docker 部署全流程实战

手把手教你用 Docker + Ray + vLLM 在多节点 GPU 集群上部署 70B+ 大模型推理服务,覆盖在线/离线安装、模型选型、性能调优与生产运维。
AI & LLM

New-API 部署同一模型双进程双端口:实现负载均衡与高可用的完整指南

Deploy two New-API processes for the same model on separate ports to achieve load balancing and high availability, with complete Docker Compose configuration and tuning notes.
AI & LLM

Running DeepSeek-V4-Flash-0731 on 8x RTX 4090D with Docker

Step-by-step guide to building a Docker image and serving DeepSeek-V4-Flash-0731 on 8x RTX 4090D consumer GPUs using the vLLM SM89 fork.
AI & LLM

Multi-Node LLM Serving: vLLM + Ray

End-to-end offline deployment of vLLM with Ray across two nodes in Docker, covering worker discovery, tensor-parallel inference, and automated health checks.
AI & LLM

Multi-Node LLM Serving: Architecture, Frameworks & Best Practices

AI-assisted
Overview of multi-node LLM serving architectures comparing vLLM, TensorRT-LLM, and SGLang, with deployment strategies for 70B+ models across GPU clusters.
AI & LLM

DeepAgents 接入 DeepSeek 配置指南

Guide for integrating DeepSeek as the LLM backend for DeepAgents, covering uv package management, API key setup, and Windows verification.