小象 v0.2
Reading intent from a turn's words alone
象 (xiàng, elephant) reads the turns I type to software agents and labels each fragment with what it is doing: proposing, explaining, approving, reporting a problem, and so on. 小象 (little elephant) is a classical model that tries to recover those labels from the words alone, with no language model and no database, so that it can run on anyone's logs.
It is deliberately simple: naive Bayes over words and word pairs, plus a short hand-written list of general English speech-act cues (below). It is expected to do poorly, and the point is to measure where. Try a sentence:
How well it does
Trained and tested on 2310 labelled fragments from 674 of my turns, with 26 intents. Each test holds out whole turns (5-fold cross-validation), so no fragment is scored by a model that saw its neighbours.
- Accuracy: 28.8%, against 15% for always guessing propose, and 28.3% without the hand-written cues.
- Macro-averaged F1 over intents: 0.12. Rare intents are mostly never predicted.
The labels are 象's own readings, and none has yet been confirmed by me. So some errors below are the model's, and some are disagreement between labellers.
Per intent
| intent | fragments | recall | precision |
|---|---|---|---|
| propose | 351 | 0.83 | 0.23 |
| explain | 269 | 0.35 | 0.21 |
| report-problem | 197 | 0.34 | 0.41 |
| approve | 188 | 0.49 | 0.73 |
| constrain | 188 | 0.16 | 0.33 |
| clarify | 181 | 0.04 | 0.30 |
| ask-action | 167 | 0.25 | 0.39 |
| report | 138 | 0.11 | 0.47 |
| qualify | 133 | 0.00 | 0.00 |
| disagree | 79 | 0.13 | 0.45 |
| redirect | 69 | 0.03 | 0.67 |
| delegate | 64 | 0.02 | 0.20 |
| verify | 61 | 0.03 | 0.40 |
| extend | 48 | 0.08 | 1.00 |
| defer | 46 | 0.00 | 0.00 |
| continue | 41 | 0.12 | 0.83 |
| prioritize | 40 | 0.05 | 0.67 |
| collect | 26 | 0.00 | 0.00 |
| retract | 14 | 0.00 | 0.00 |
| ask-information | 2 | 0.00 | 0.00 |
| challenge | 2 | 0.00 | 0.00 |
| provide-reference | 2 | 0.00 | 0.00 |
| concede | 1 | 0.00 | 0.00 |
| contrast | 1 | 0.00 | 0.00 |
| reframe | 1 | 0.00 | 0.00 |
| withdraw | 1 | 0.00 | 0.00 |
Most common mistakes
| labelled | predicted | count |
|---|---|---|
| explain | propose | 141 |
| constrain | propose | 115 |
| ask-action | propose | 95 |
| clarify | propose | 90 |
| report-problem | propose | 72 |
| qualify | propose | 71 |
| report | propose | 61 |
| clarify | explain | 59 |
| approve | propose | 50 |
| report-problem | explain | 42 |
Intent collisions
The same short text labelled with different intents. No classifier that reads only the words can get all of these right; they need context.
| text | labels |
|---|---|
| ok | approve 11, continue 2 |
| please continue | continue 2, ask-action 1 |
| hm | report 1, continue 1, qualify 1 |
| just say hi | ask-action 1, verify 1 |
| quote | collect 1, explain 1 |
Hand-written cues
Words whose force is the same in any conversation, added as pseudo-counts (5 each) rather than learned. Version 0.1 read “I reject your claim” as approval: “reject” occurs in only four of my fragments, three of them about papers or ideas being rejected, so it never entered the model, and the decision rested on “I” and “your”.
| intent | cues |
|---|---|
| disagree | i reject, reject, i disagree, disagree, that's wrong, wrong, incorrect, i object, not right |
| approve | i agree, agree, sounds good, looks good, that's right, approved, go ahead, great |
| report-problem | doesn't work, not working, broken, failed, fails, bug |
| ask-action | could you, can you, would you |
When a sentence contains only very common words, the page says so instead of guessing.
Model: 2476 tokens, each seen in at least 3 separate turns, with identifiers, names of people and places, and anything a secret scanner flags removed. Built 2026-09-29.