GPU / CUDA / Docker 部署兼容性清单:从镜像到模型服务
用版本矩阵、镜像 digest、模型校验、GPU 可见性、安全边界和功能验收,把消费级 GPU 部署记录整理成可复用的兼容性案例。
从模型调用到 Agent 系统:检索、工作流、推理服务、工具权限与生产化实践。
把 RAG、LangChain、LangGraph、DeepAgents、OpenCode 和模型网关等历史文章整理成一条从 API 调用到可观测 Agent 系统的学习路线。
Read note →把 Ray wheelhouse、GPU 容器、并行策略、私有网络、认证、健康检查和回滚拆开,避免把一条历史启动命令误当成生产保证。
Read note →整合 vLLM、Ray、Docker、模型网关和离线部署笔记,给出模型服务从容量估算到灰度、监控与回滚的生产检查表。
Read note →用版本矩阵、镜像 digest、模型校验、GPU 可见性、安全边界和功能验收,把消费级 GPU 部署记录整理成可复用的兼容性案例。
Read note →以威胁模型为中心整理 DeepAgents、OpenCode、Agent Swarm 和文件服务实践中的最小权限、网络、密钥、审批、预算与审计设计。
Read note →用版本矩阵、镜像 digest、模型校验、GPU 可见性、安全边界和功能验收,把消费级 GPU 部署记录整理成可复用的兼容性案例。
把 Ray wheelhouse、GPU 容器、并行策略、私有网络、认证、健康检查和回滚拆开,避免把一条历史启动命令误当成生产保证。
以威胁模型为中心整理 DeepAgents、OpenCode、Agent Swarm 和文件服务实践中的最小权限、网络、密钥、审批、预算与审计设计。
整合 vLLM、Ray、Docker、模型网关和离线部署笔记,给出模型服务从容量估算到灰度、监控与回滚的生产检查表。
把 RAG、LangChain、LangGraph、DeepAgents、OpenCode 和模型网关等历史文章整理成一条从 API 调用到可观测 Agent 系统的学习路线。
Introduction to Agent Swarm, an open-source multi-agent AI operating system, with a step-by-step focus on offline deployment of its Docker-based Lead and Worker architecture.
Learning roadmap for LLM application development, mapping RAG, LangChain, LangGraph, and DeepAgents from fundamentals to production-grade agent workflows.
手把手教你用 Docker + Ray + vLLM 在多节点 GPU 集群上部署 70B+ 大模型推理服务,覆盖在线/离线安装、模型选型、性能调优与生产运维。
Deploy two New-API processes for the same model on separate ports to achieve load balancing and high availability, with complete Docker Compose configuration and tuning notes.
OpenCode 是一个开源的 AI 编程助手,运行在终端中。本文深入解析其 TUI 架构、声明式 UI 渲染、Client-Server 分离设计,并手把手教你用 Bun + SolidJS + OpenTUI 构建一个类似的 AI CLI 工具。
Step-by-step guide to building a Docker image and serving DeepSeek-V4-Flash-0731 on 8x RTX 4090D consumer GPUs using the vLLM SM89 fork.
End-to-end offline deployment of vLLM with Ray across two nodes in Docker, covering worker discovery, tensor-parallel inference, and automated health checks.
Overview of multi-node LLM serving architectures comparing vLLM, TensorRT-LLM, and SGLang, with deployment strategies for 70B+ models across GPU clusters.
Overview of MiniCPM5-1B, a 1B-parameter on-device language model from 面壁智能, with benchmarks, architecture details, and deployment considerations.
Offline deployment of MiniCPM5-1B using llama.cpp server in Docker, with Open-WebUI frontend and CPU-only inference configuration.
Offline deployment of MiniCPM5-1B using Ollama and Open-WebUI in Docker, covering image loading, model import, and CPU-only inference setup.
Setup and usage notes for Yuxi, an open-source knowledge base and knowledge graph agent platform for building searchable document collections.
Overview of OpenCode, an open-source AI coding assistant with intelligent completion, multi-model support, and a CLI-based conversational workflow.
Self-hosted gateway for unifying access to multiple AI providers, with usage tracking, key management, and OpenAI-compatible API endpoints.
Notes on using Sentence Transformers for semantic search and text similarity, covering embedding generation, reranking, and integration with vector stores.
Guide for integrating DeepSeek as the LLM backend for DeepAgents, covering uv package management, API key setup, and Windows verification.
Private deployment of DeepAgents using sandbox backends for isolated code execution, with Docker-based sandboxes and custom backend configuration.