LOADING

Archive

A chronological list of all published posts.

2026
48 posts
08-16
Video Audio Multimodal —— 视频理解 & 视频生成 & 音频处理
08-14
Vision-Language-Action —— VLA具身智能
08-12
Multimodal RAG —— 多模态RAG
08-10
Denoising Diffusion Probabilistic Models —— DDPM扩散模型
08-08
Multimodal Large Language Model —— MLLM多模态大语言模型
08-06
Vision-Language Model —— VLM视觉语言模型
08-04
Vision Transformer —— ViT视觉Transformer
08-02
Contrastive Language-Image Pre-training —— CLIP对比学习
07-16
RLVR (Reinforcement Learning with Verifiable Rewards) —— 可验证奖励强化学习
07-14
大模型强化学习对齐方案:PPO、GRPO、DPO、DAPO与GSPO技术详解与对比
07-12
GSPO (Group Sequence Policy Optimization) -- 群组序列策略优化
07-10
DAPO (Decoupled Clip and Dynamic sAmpling Policy Optimization) -- 解耦裁剪与动态采样策略优化
07-08
GRPO (Group Relative Policy Optimization) -- 群体相对策略优化
07-06
DPO (Direct Preference Optimization) -- 直接偏好优化
07-04
PPO(Proximal Policy Optimization) -- 近端策略优化
07-02
强化学习基础 -- 策略梯度、优势函数、重要性采样与KL散度惩罚
06-30
RLHF (Reinforcement Learning from Human Feedback) —— 基于人类反馈的强化学习
06-28
SFT (Supervised Fine-Tuning)—— 监督微调
06-26
PEFT (Parameter-Efficient Fine-Tuning) —— 参数高效微调
04-18
MLOps & Observability
04-16
TTFT & TPOT & Throughput —— LLM推理性能指标
04-14
Quantization & Flash Attention —— LLM推理优化技术(二)
04-12
Speculative Decoding —— LLM推理优化技术(一)
04-10
Continuous Batching & Chunked Prefill —— LLM推理引擎调度
04-08
PagedAttention & RadixAttention —— LLM推理引擎架构
04-06
NCCL & DP/MP/PP/FSDP/ZeRO —— 分布式训练
04-04
Shared Memory & Coalesced Access —— CUDA性能优化技术
04-02
GPU & Grid-Block-Thread —— CUDA编程心智模型
04-01
vLLM与SGLang:大模型推理框架的全面对比
04-01
SGLang:大模型推理的"结构化程序"革命
04-01
vLLM:大模型推理系统的分页内存革命
03-16
Multi-Agent System—— 多智能体系统
03-14
Agent Communication Protocols —— MCP & A2A
03-12
Agent Engineering —— Prompt & Context & Harness & Loop
03-10
Memory System —— Agent记忆系统
03-08
Retrieval-Augmented Generation —— 检索增强生成
03-06
Function Calling —— 工具调用
03-04
Agent Paradigms —— ReAct & Plan-and-Solve & Reflection
03-02
Agent Architecture —— LLM & Planning & Tool & Memory
02-18
Lecture 9 —— LLM多模态泛化与前沿探索:从ViT/VLM跨模态适配到扩散LLMs与未来发展趋势
02-16
Lecture 8 —— LLM评估全解:从人工评分、LLM-as-a-Judge到基准测试与工具调用故障排查
02-14
Lecture 7 —— Agentic LLMs全解析:从RAG检索增强、工具调用(Tool Calling)到多智能体(Agent)协作
02-12
Lecture 6 —— LLM推理能力深度解析:从思维链(CoT)到GRPO强化学习与DeepSeek R1训练范式
02-10
Lecture 5 —— LLM偏好调优全解:从RLHF人类反馈强化学习到DPO直接偏好优化
02-08
Lecture 4 —— LLM训练全流程解析:从预训练缩放规律到参数高效微调与模型对齐
02-06
Lecture 3 —— 大语言模型(LLM)核心技术与推理优化全解:从MoE架构到高效解码
02-04
Lecture 2 —— Transformer进阶:位置编码演进、注意力优化技巧与BERT预训练微调全解
02-02
Lecture 1 —— 从RNN到Transformer:NLP序列建模演进与编码器-解码器架构全解
2025
25 posts
08-22
加密文档测试
08-20
LSTM & GRU —— 门控循环神经网络
08-18
RNN & BPTT —— 循环神经网络与随时间反向传播
08-16
Residual Network —— ResNet残差网络
08-14
CNN —— 卷积神经网络
08-12
Activation & Initialization —— 激活函数与权重初始化
08-10
MLP & Back Propagation —— 多层感知机与反向传播
08-08
Autograd & Computational Graph —— 自动微分与计算图
08-06
Gradient Descent Optimizer —— 梯度下降优化器
08-04
Probability & Information —— 概率论与信息论基础
08-02
Tensor Operations —— 向量、矩阵与张量运算
07-22
K-Means Clustering —— 聚类算法
07-20
XGBoost & LightGBM —— 梯度提升框架
07-18
Gradient Boosting Machine (GBM) —— 梯度提升机与加法模型
07-16
Bagging & Random Forest —— 随机森林
07-14
Decision Tree —— 决策树
07-12
Naive Bayes —— 朴素贝叶斯
07-10
K-Nearest Neighbor (KNN) —— K-近邻
07-08
Kernel Trick —— 核技巧与常用核函数
07-06
Support Vector Machine (SVM) —— 支持向量机
07-04
Logistic Regression —— 逻辑回归与Softmax多分类
07-02
Linear Regression —— 线性回归
06-30
Docker 快速配置深度学习环境 + 基础命令 + 常见报错
06-28
开发基础工具 —— Conda、Pip/UV 与 Git 实战手册
06-26
命令行基础指令 —— Linux Shell 与 APT 包管理速查
Directory
Albums
Diary
Posts
Projects
Skills
Timeline
Categories
Tags
Table of Contents
Statistics