介绍

前言

前段时间 TypeSafe AI JEV 凭借快速决策能力出圈,斯坦福 + Hazy+NVIDIA 推出 CLM8b 也具有同等能力:

两者定位完全一致:不做文本生成,专门做 Agent 快速决策、候选动作打分、选择、yes/no 判断;接口原语对齐:Choice / Score / Noul。不同点如下:

部署

环境

系统

Ubuntu 24.04.4 LTS + NVIDIA GeForce RTX™ 4070 Laptop GPU,可以使用 sudo ubuntu-drivers install 补充缺少的驱动
Nvidia驱动

Python

使用 conda,安装 pytyhon=3.13.16

代码

github

https://github.com/Contrastive-LM/CLM,下载完成后安装所有依赖\

启动

前置准备,获取相应模型(如果遇到无法获取模型可先进行设置镜像:export HF_ENDPOINT=https://hf-mirror.com):\

# 1. encoder (Qwen3-8B embeddings)
vllm serve Qwen/Qwen3-8B --served-model-name qwen3-8b --runner pooling --max-model-len 2048 --port 8090 --gpu-memory-utilization 0.9

第一步需要下载模型作为编码作用,将输入的状态转换成数字提供给后续使用
vLLM 启动成功
vLLM启动成功

# 2. CLM API on :8700 (downloads the 75 MB reference head on first run)
clm-serve

第二步启动页面交互,将状态通过向量匹配得到结果
CLM 启动成功
CLM启动成功

结构图
结构图

  1. 骨干底座:冻结 Qwen3‑8B(底座权重完全不训练,只做编码)

  2. 两个可训练投影头:State Head(状态编码器)、Action Head(动作编码器),合计仅 20M 参数(75MB checkpoint)

  3. 训练目标:双向 InfoNCE 对比损失

- 把 “正确状态‑动作” 向量拉近;错误样本向量推开

  1. 推理逻辑:

- 分别编码当前状态s、所有候选动作a₁,a₂…得到 embedding

- 计算向量点积 / 余弦相似度打分,softmax 输出分布,选最高分动作

  1. 核心特性:embedding 可独立缓存

如果候选动作集合不变(工具列表、菜单选项),动作向量只算一次反复复用;只需要重新编码变化的状态,极大降低 Agent 循环延迟。

使用

页面使用

启动成功之后可通过浏览器访问:http://localhost:8700
Choice / Score / Noul
Choice / Score / Noul

Rank
Rank

以上为 CLM8b 两种使用方法,Rank 目前不够稳定

接口使用

CURL 请求

curl -s http://localhost:8700/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
  "state": "Customer: my invoice was charged twice and nobody answers the phone!",
  "model": "clm-latest",
  "questions": {
    "urgency": {
      "type": "noul",
      "instructions": "Is this urgent?"
    },
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "billing": "Charges, invoices, refunds",
        "technical": "Bugs and outages"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": [
        "Calm",
        "Frustrated",
        "Very angry"
      ]
    }
  }
}'

Response 返回

{
        "model": "clm-latest",
        "answers": {
                "urgency": {
                        "type": "noul",
                        "noul": 0.8263707919505312
                },
                "department": {
                        "type": "choice",
                        "choice": "billing",
                        "confidence": 0.9871286986216236,
                        "probabilities": {
                                "billing": 0.9935643493108117,
                                "technical": 0.00643565068918816
                        }
                },
                "frustration": {
                        "type": "score",
                        "score": 1.9999515815375033,
                        "confidence": 0.9999452154944316,
                        "legend": {
                                "0": "Calm",
                                "1": "Frustrated",
                                "2": "Very angry"
                        },
                        "probabilities": {
                                "0": 1.1895458784263167e-05,
                                "1": 2.4627544927911903e-05,
                                "2": 0.9999634769962877
                        }
                }
        },
        "usage": {
                "billing_units": 3,
                "input_tokens": 0,
                "output_tokens": 0
        }
}


↙↙↙阅读原文可查看相关链接,并与作者交流