clef-text-0.6b

A 0.70B-parameter text decision model distilled from Cloudflare's clef-flash (9B): Qwen3-0.6B (LoRA merged) + Clef's joint schema head. Same Jev / SystemOne contract: one probability per option of every typed question, one forward pass. Good at routing and classification; not a knowledge model (see the table).

Apple silicon / Neural Engine build: FluidInference/clef-text-0.6b-coreml.

Usage

joint_schema_model.py is Cloudflare's unchanged Clef module (record encoding, head, ClefModel).

import json, sys, torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from transformers import AutoModelForCausalLM, AutoTokenizer

path = snapshot_download("FluidInference/clef-text-0.6b")
sys.path.insert(0, path)
from joint_schema_model import ClefModel, JointSchemaHead, collate_records, encode_record

tokenizer = AutoTokenizer.from_pretrained(path)
backbone = AutoModelForCausalLM.from_pretrained(path, dtype=torch.float32)
head = JointSchemaHead(**json.loads(open(f"{path}/joint_head_config.json").read()))
head.load_state_dict(load_file(f"{path}/joint_head.safetensors"))
model = ClefModel(backbone, head).eval()

record = {"state": "Our checkout is down and customers can't pay.",
          "questions": {"team": {"type": "choice", "instructions": "Which team should handle this?",
                                 "criteria": {"billing": "Payments", "engineering": "Bugs or outages"}},
                        "urgency": {"type": "score", "criteria": ["Low", "Normal", "High", "Critical"]}}}
encoded = encode_record(tokenizer, record)
with torch.inference_mode():
    logits = model(collate_records([encoded], tokenizer.pad_token_id, torch.device("cpu")))[0]
for question, question_logits in zip(encoded.questions, logits):
    print(question.question_id, dict(zip(question.option_ids, question_logits.softmax(-1).tolist())))

Quality (held-out test splits, gold accuracy)

Task this model clef-flash 9B
DBpedia-14 99.7 100.0
AG News 90.7 92.0
SST-2 90.0 94.0
BANKING77 86.4 94.5
TweetEval offensive 85.0 82.3
Yelp stars 69.0 70.7
Emotion 68.0 58.0
BoolQ 82.3 90.7
MNLI 72.3 84.3
ARC-Easy 80.0 100.0
CommonsenseQA 65.0 94.3
ARC-Challenge 60.7 97.3
all (5,124 q) 79.1 88.2

Training: 24,208 records from 12 public datasets (train splits) labelled by clef-flash, KL (T = 2) + 0.3 gold CE, LoRA r64 + full head, 2 epochs on an M5 Pro.

License and credits

Apache-2.0. joint_schema_model.py and LICENSE are Cloudflare's (Apache-2.0). Teacher Cloudflare/clef-flash; base Qwen/Qwen3-0.6B. By Fluid Inference.

Downloads last month
21
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/clef-text-0.6b

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1349)
this model
Quantizations
1 model