Model Articles
10 articles
GPU / CUDA / Docker 部署兼容性清单:从镜像到模型服务
Curated guide用版本矩阵、镜像 digest、模型校验、GPU 可见性、安全边界和功能验收,把消费级 GPU 部署记录整理成可复用的兼容性案例。
New-API 部署同一模型双进程双端口:实现负载均衡与高可用的完整指南
Deploy two New-API processes for the same model on separate ports to achieve load balancing and high availability, with complete Docker Compose configuration and tuning notes.
Running DeepSeek-V4-Flash-0731 on 8x RTX 4090D with Docker
Step-by-step guide to building a Docker image and serving DeepSeek-V4-Flash-0731 on 8x RTX 4090D consumer GPUs using the vLLM SM89 fork.
Multi-Node LLM Serving: vLLM + Ray
End-to-end offline deployment of vLLM with Ray across two nodes in Docker, covering worker discovery, tensor-parallel inference, and automated health checks.
Multi-Node LLM Serving: Architecture, Frameworks & Best Practices
AI-assistedOverview of multi-node LLM serving architectures comparing vLLM, TensorRT-LLM, and SGLang, with deployment strategies for 70B+ models across GPU clusters.
MiniCPM5-1B Overview
Overview of MiniCPM5-1B, a 1B-parameter on-device language model from 面壁智能, with benchmarks, architecture details, and deployment considerations.
Deploying MiniCPM5-1B with llama.cpp
Offline deployment of MiniCPM5-1B using llama.cpp server in Docker, with Open-WebUI frontend and CPU-only inference configuration.
Deploying MiniCPM5-1B with Ollama
Offline deployment of MiniCPM5-1B using Ollama and Open-WebUI in Docker, covering image loading, model import, and CPU-only inference setup.
AI Model Hub: New API
Self-hosted gateway for unifying access to multiple AI providers, with usage tracking, key management, and OpenAI-compatible API endpoints.
Sentence Transformers: Sentence-BERT
Notes on using Sentence Transformers for semantic search and text similarity, covering embedding generation, reranking, and integration with vector stores.
Tag navigation
Featured tags