If you want an early read on which claim conversations are heading toward an attorney, the language on both sides of the exchange is a good place to start. Using 5,000 synthetic claim chats, this walk-through shows you how to count LIWC-style categories for each speaker, keep conversation length from fooling you, predict from the first two turns with a calibrated model, check a proxy against a thin survey anchor, and test whether it holds up in a second population.
When a service conversation starts to fail, the person helping tends to say so first, in apologies and stated limits, while most dissatisfied customers leave without a word of complaint; anyone who waits for the customer to object starts late and misses most of the cases they set out to find. I first built a version of this method for a conversational assistant answering customer requests, and the part that stayed with me was how early the helper’s side of the transcript started apologizing. Here the method runs end to end on 5,000 synthetic claim chats between claimants and claim professionals, with attorney representation within 90 days as the outcome, so every step can be shown on data anyone may inspect. You will see seven dictionary categories in the LIWC tradition, three for the professional and four for the claimant; how a whole-conversation flag picks up the outcome through chat length, and how a rate inside a fixed first-contact window closes that route; a first-two-turn model that predicts representation (AUC .65, calibration slope .96); a weak-supervision label model built from three evidence streams and measured against 300 survey responses, 20 of them negative; and the same pipeline rerun on a second synthetic population, where the proxy falls to chance. The method ends in a pre-registered field experiment, because an indicator earns its keep once it points at something a claim organization can test.
Satisfied people leave very little text behind. A claimant who gets a clear answer about the rental, the inspection date, or the payment tends to say thanks and close the chat, and the transcript of a conversation that went well is short, polite, and nearly empty of anything a dictionary could count. Dissatisfaction has two routes, and Hirschman (1970) named them for customers of declining firms; a dissatisfied person can use voice, which in a chat looks like a correction or an angry line, or exit, which looks like nothing at all. In conversation logs most dissatisfied people choose exit, so a team that waits for the customer to complain starts late and misses most of the cases it set out to find.
The person on the other side of the exchange behaves differently. A claim professional who cannot approve an amount, cannot move a date, or has to explain a gap in coverage says so in words, and in those words the breakdown is visible in real time; “I’m sorry for the wait,” “unfortunately I can’t approve that today,” “I think the shop might have it by then.” The helper narrates the failure as it happens, whether or not the claimant ever answers back. Researchers who studied malpractice plaintiffs found a parallel pattern at the far end of the process; families who sued after a perinatal injury often cited communication problems, such as feeling misled or unable to get information, among their reasons for seeking a lawyer (Hickson et al., 1992), and plaintiffs in depositions repeatedly named relationship problems, such as feeling deserted or unheard, alongside the clinical facts (Beckman et al., 1994).
In claim work, people sit on both sides of the chat and the outcome at the end is costly and observable, which makes claims a natural home for this kind of measurement, and a synthetic population makes the transposition concrete. In 5,000 simulated chats with a representation rate of 11.6 percent, the 29 percent of chats in which the professional apologized at least once ran 7.8 turns on average against 4.4 for the rest (1.8 times as long), closed in a single exchange 13 percent of the time against 51 percent, and ended in representation 26 percent of the time against 5.6 percent. That contrast is reported here the way a dashboard usually reports it, as any apology anywhere in the chat, and section 3 takes part of it back.
The counting method comes from the Linguistic Inquiry and Word Count (LIWC) tradition, in which researchers assign words to psychologically meaningful categories and report each category as a percentage of the words in a text (Pennebaker et al., 2015; Tausczik & Pennebaker, 2010). The finding from that tradition most relevant here is that much of the social and psychological information in a text sits in its small, frequent words; pronouns, negations, and other function words vary with attention, self-focus, and stance in ways speakers seldom monitor (Chung & Pennebaker, 2007; Pennebaker et al., 2003). First-person singular is the clearest case. Higher rates of “I,” “me,” and “my” accompany distress and depression, modestly but consistently across many samples (Rude et al., 2004; Tackman et al., 2019), and in conversations between people of unequal standing the lower-status speaker tends to use more of them (Kacewicz et al., 2014), a description that fits a claimant waiting on a decision from the person across the chat.
Seven categories do the work here, three for the claim professional and four for the claimant, and each one is counted only in its own speaker’s turns, with that speaker’s own word count as the denominator. The same word means different things on each side of the exchange; “can’t” from a professional states a limit on what the company will do, while “can’t” from a claimant is negation, often about their own situation (“I can’t keep paying for this myself”). Four categories have close LIWC2015 analogues. Apology, limitation, and correction are specific to service conversations, so they are custom lists, and LIWC users can load lists like these as user dictionaries. The lexicon below is a small illustrative one of my own; for a real deployment it helps to take the validated LIWC-22 categories where they exist (Boyd et al., 2022) and to validate the custom lists against human coding before counting anything at scale.
| Category | Speaker | Nearest LIWC2015 category | Example entries | Role in the simulation |
|---|---|---|---|---|
| Apology | Claim professional | none (custom) | sorry, apologize, apologies, regret | raises representation risk |
| Limitation | Claim professional | none (custom) | unfortunately, unable, cannot, can’t, not able, not authorized | raises representation risk |
| Hedging | Claim professional | tentat | maybe, perhaps, possibly, might, probably, I think, I believe | planted null |
| Negation | Claimant | negate | no, not, never, nothing, nobody, any word ending in n’t | raises risk modestly |
| Correction | Claimant | none (custom) | already, I said, I told, I meant, actually, twice | raises risk modestly, arrives late |
| Anger | Claimant | anger | angry, furious, ridiculous, unacceptable, fed up, sick of | raises risk modestly, arrives late |
| First-person singular | Claimant | i | I, me, my, mine, myself, I’m, I’ve | raises risk modestly |
In the last column sits the one thing you can never know in a real dataset, the truth. I generated every chat from fixed sentence banks under a planted process, with a latent friction score for each chat. Higher friction means more professional apologies and limitations, slightly more claimant negation and first-person singular, more correction and anger only for the 30 percent of claimants who voice complaints at all (and more so in later turns), and a higher probability of representation. Hedging I planted at a constant rate per sentence, unrelated to friction or outcome, so hedging is a null category; a method that finds hedging predictive has found an artifact. Because the truth is known, every analytic step can be scored against it, which is the whole reason to build the example synthetically.
I feel like I'm on my own with this.
Sorry about the mix-up with the rental. Unfortunately the policy does not cover that item.
A tree from next door fell on the garage. The adjuster came out to look at the roof on Monday.
I regret that the callback took this long. I think the shop might have it by then. My apologies for the delay on the payment.
I don't know what I'm supposed to do now. Evenings are better for a call. I can't keep paying for this myself.
Hopefully we can close this out soon. I'm sorry, that request was missed on our end.
That's not what I said earlier. I uploaded the photos of the damage this morning. What do you need from me to move this forward?
The easiest way to turn a dictionary into a conversation-level measure is a flag; a chat “has apology” if any professional turn contains an apology word. The flag is intuitive, and its failure is predictable. If each sentence has a small probability p of containing a category word, a chat with n sentences contains at least one with probability 1 minus (1 minus p) to the power n, which climbs steadily with n. If some chats run longer than others for any reason, every flag moves with them, whatever the words mean.
In the simulation, length depends partly on the outcome itself; claimants headed toward representation ask more questions and stay in the chat longer, so turn count correlates r = .42 with representation, and part of that correlation passes into every whole-chat flag. The figure below has a toggle between the two ways of counting. Counted as any occurrence anywhere in the chat, hedging correlates r = .16 (95% CI .13 to .18) with representation; counted as a rate per 100 words in the first two turns, the same category correlates r = .01 (CI −.02 to .04), consistent with the zero I planted. With log turn count held constant, the whole-chat hedging flag has a partial r of .04, so nearly all of its apparent signal was length.
Switching to a rate in a fixed window removes two things at once, and it helps to keep them apart. The first is length leakage, which the hedging row isolates. The second is timing; claimant correction and anger have real planted signal, but they arrive late in the chat, so little of them shows in the first-contact window (correction falls from r = .11 as a whole-chat flag to r = .00 as an early rate), and their whole-chat flags shrink to partial correlations of .02 and .03 once length is held constant. Apology loses the least, from r = .29 as a flag to r = .18 as an early rate, because its signal was planted in every professional turn, the first reply included.
LIWC researchers have long handled the length problem by reporting each category as a percentage of words (Pennebaker et al., 2015); the trap reappears whenever a team converts rates back into flags for a dashboard, which happens more often than anyone would like to admit. The cost of the trap is easiest to see in a model. A logistic regression on the seven whole-chat flags reaches an AUC of .79 on the same synthetic data, well above the first-contact model in the next section; most of that margin is chat length, which a claim organization observes only after the conversation is over and which, here, depends on the outcome itself.
At first contact only the first two turns exist, so the features are the seven category rates in turns 1 and 2, the claimant’s opening message (14.7 words on average) and the professional’s first reply (17.9 words). The model is a logistic regression on standardized rates with a light L2 penalty, scored by five-fold stratified cross-validation so every chat is scored by a model fit without it, with a 95% interval for the AUC from 2,000 bootstrap resamples of those out-of-fold predictions (Efron & Tibshirani, 1993; Hanley & McNeil, 1982). Resampling the evaluation set alone leaves out the smaller variation from refitting the model, which is worth knowing when you read the interval.
The cross-validated AUC was .65 (95% CI .62 to .67), a modest level of discrimination, and in my experience the level a first-message model tends to reach on real service conversations as well.
Discrimination is how well the model ranks claims; calibration is whether its probabilities mean what they say, and a measure read as rates by team or by week is only as good as its calibration (Van Calster et al., 2019). The model is fit without class reweighting, since with a reweighted minority class the AUC stays the same and every predicted probability rises. Mean predicted risk was 11.6 percent against an observed rate of 11.6 percent, the calibration slope was .96 (1.00 is ideal), and the Brier score was .098 against .102 for predicting the base rate for everyone. The top decile of predicted risk had an observed representation rate of 28.8 percent and the bottom decile 9.2 percent.
The coefficients recover the planted structure, with one instructive exception.
| First-two-turn category | Odds ratio per SD | 95% bootstrap CI |
|---|---|---|
| Apology (professional) | 1.50 | 1.40 to 1.61 |
| Limitation (professional) | 1.33 | 1.23 to 1.42 |
| Hedging (professional) | 1.11 | 1.01 to 1.21 |
| Negation (claimant) | 1.13 | 1.03 to 1.21 |
| Correction (claimant) | 0.98 | 0.86 to 1.05 |
| Anger (claimant) | 1.06 | 0.97 to 1.12 |
| First-person singular (claimant) | 1.04 | 0.95 to 1.14 |
Apology and limitation have the largest odds ratios. Negation has the modest effect I planted; correction and anger sit near 1.0 because they seldom appear in the first claimant turn; and first-person singular, planted at the same strength as negation, loses its signal in an opening message of about fifteen words, a reminder that dictionary rates on very short texts are noisy. Hedging, the planted null, comes out at 1.11 with an interval that just clears 1.0, and the cause is a feature of the generator. Each sentence has at most one category, so a hedge occupies a slot an apology could have filled; once apology is held constant in the regression, the hedging coefficient rises above one, a suppression effect created by the way the text was built. The same kind of artifact appears in real text whenever categories compete for the same words or slots, and the univariate rate (r = .01) is the safer reading of hedging here.
The timing of each signal shows the same asymmetry. In the synthetic population, 43 percent of all chats had a professional apology or limitation somewhere, against 8.6 percent with a claimant correction, so the helper’s side flags 5.0 times as many chats as the claimant’s. The ratio in this simulation is held down by the generous assumption that 30 percent of claimants voice complaints; real customers tend to voice even less, so expect the ratio on your own logs to run higher.
Among the 578 chats that ended in representation, 79 percent had a helper signal by the end of the chat and 19 percent had a claimant correction or anger, and 81 percent of those claimants never voiced either. The helper signal also arrived first; it appeared in the second turn for 47 percent of represented chats, and the median first helper signal fell at turn 2 against turn 5 for the claimant. By turn 3 only 6.4 percent of represented chats contained a claimant correction or anger.
The other half of the lesson is in the dashed lines. By the end of the chat, 39 percent of the chats that did not end in representation also had a helper signal, so the helper’s language is sensitive and far from specific; an apology marks friction, and most friction resolves. The signal belongs in a rate across many claims (a team, a line of business, a week), where its sensitivity is a strength and its lack of specificity averages out.
Representation is a hard outcome, and claim professionals record it in the file. The claimant’s experience is softer and matters on its own terms, and here the only explicit evidence is a post-close survey that few people answer. In the simulation, 300 survey responses were drawn at random, and 20 of them were negative, about the size of anchor a team usually has when it first asks whether its implicit measures mean anything.
Three evidence streams combine into one label. The first is helper language (the professional’s apology and limitation rate over the whole chat, flagged for the top fifth); the second is claimant language (negation, correction, and anger, flagged the same way); the third is behavioral (a repeat contact within seven days). Rates per word keep the length trap out of the two language streams, and reading the whole chat is appropriate here, since the question concerns how the conversation went once it was over. A two-class latent class model fit by expectation maximization (Dawid & Skene, 1979), the model family behind data programming and weak supervision (Ratner et al., 2016, 2017), estimates each stream’s sensitivity and false-positive rate without any labels and yields a posterior probability of a dissatisfied experience for every chat. A simulated LLM-judge label stays out of the model and serves as an independent cross-check, the way a judge pass is often used to audit a proxy (Zheng et al., 2023).
Against the 300-response anchor, the label model’s AUC is .64 (95% CI .51 to .77). In the simulation you can do something no real program allows and score the same label against every chat’s true state; that AUC was .76. Drawing 1,000 alternative surveys of 300 from the same population, the middle 90 percent of their AUC estimates ran from .61 to .82. The anchor that happened to be drawn sits low in that spread, and with any single anchor of this size the answer would stay just as open; by the Hanley and McNeil (1982) approximation the AUC on the current anchor is known only to within about ±.13, and narrowing that to ±.05 at this prevalence takes about 2,190 responses, roughly 146 of them negative.
Fitting the label model means assuming the streams are independent once the true state is known, and dropping each stream in turn is a practical check on that assumption. Without helper language, the true AUC fell from .76 to .71; without repeat contact it fell to .70; without claimant language it rose to .78. Both language streams rise with the same friction, and among the satisfied survey respondents their flags still correlated (phi = .19), so the shared friction enters the posterior twice and gets too much weight. None of this was visible on the anchor, since its estimates for the four versions (.64, .63, .66, .62) differ by far less than its own interval. The judge’s labels and the proxy agreed only weakly, at a Cohen’s (1960) kappa of .16, although the judge’s own AUC against the true state was .71; two imperfect readers of the same construct can disagree substantially while each contains real signal.
Anyone who validates an index in one population and uses it in another makes an implicit claim that the process generating its signals stays the same, and the claim is worth testing on purpose. The test here is a second synthetic population of 5,000 chats built on a different process, the kind of difference that separates two lines of business, or one team before and after a new script. Every professional in Population B opens with a scripted apology, apology and limitation language appear at constant rates whatever the friction, nearly every claim receives a scheduled callback (97 percent have a repeat contact), and dissatisfaction and representation depend mainly on something absent from every transcript, the timing of a payment after the chat ends. The identical pipeline runs again.
| Measure | Population A | Population B |
|---|---|---|
| Label-model AUC on the 300-response anchor | .64 (.51 to .77) | .40 (.32 to .50) |
| Label-model AUC against every chat’s true state | .76 | .50 |
| Middle 90% of 1,000 alternative anchors | .61 to .82 | .42 to .59 |
| Agreement of judge and label (Cohen’s kappa) | .16 | −.01 |
| First-contact representation model, trained on A | .65 (.62 to .67) | .50 (.47 to .52) |
In B the label model’s AUC against the true state is .50, the judge and the label agree at chance, and B’s claims come out at AUC .50 under the first-contact model trained on A. The anchor estimate of .40 is a further hazard of a thin anchor, since an index with no information can land below chance by luck. Population B is built the way real failures tend to happen; a scripted opening removes the variance from the helper stream, a scheduled callback removes it from the behavioral stream, and the thing that drives the outcome never appears in the text.
The practical conclusion is to collect explicit feedback before trusting the index, and to keep collecting it. A one-question post-close survey with a reason code, sized to measure the AUC within ±.05 in each population and re-read whenever a script or process changes, costs far less than acting on an index that has stopped measuring anything.
From an association between apology language and representation you learn that apologies mark breakdowns, and nothing yet about whether a different way of handling the breakdown would change the outcome. A field experiment answers that question, and the indicator is how you choose what to test. Experimental work on apologies in legal settlement suggests that an apology that accepts responsibility and one that only expresses sympathy may move claimants differently (Robbennolt, 2003), and the limitation category points at a second target, the moment a professional tells a claimant what the company will not do. A pre-registered test, filed before any claim enters it (Nosek et al., 2018), would carry these elements.
The sample size follows from the outcome. Detecting a drop from 11.6 to 10.15 percent (a 12.5 percent relative reduction) at α = .05 with 80 percent power requires 7,236 claims per arm if claims were independent. Claims are clustered within teams, so the requirement grows by the design effect, one plus the intraclass correlation times the cluster size minus one.
| Claims per team | Assumed ICC | Design effect | Claims per arm | Teams per arm |
|---|---|---|---|---|
| 200 | .01 | 2.99 | 21,636 | 109 |
| 200 | .02 | 4.98 | 36,036 | 181 |
| 400 | .01 | 4.99 | 36,108 | 91 |
The ICC values are assumptions, and it pays to estimate them from historical claims before committing to a design, since the required sample at an ICC of .02 is about two thirds larger than the sample at .01.
The measures belong in aggregate, as rates by team, line of business, and week. With a first-contact AUC of .65, a large share of individual claims would be misranked, and a claimant scored by the words they used and then handled differently because of the score has been treated unfairly on weak evidence, in a way that could produce the very outcome the score predicted. The professional’s language is the more appropriate object of measurement in any case, since it is a record of how the organization handled the conversation, and the use it supports best is coaching and process design, applied to teams.
The associations are observational. Apology language marks a breakdown; whether a different response to the same breakdown would change representation is a causal question, and the pre-registered experiment in section 8 is the way to answer it.
A few things tend to bite in practice. Rates computed from about 33 words per chat are noisy, as the first-person singular result shows, so a very short opening window asks a lot of a dictionary. Real post-close surveys are self-selected where the simulated one was a random sample, so a real anchor needs a nonresponse analysis before it can anchor anything. The label model assumes conditional independence, and language streams that rise with the same friction tend to violate it (phi = .19 here), so it helps to run the drop-one check every time and to hold the posterior lightly where the streams overlap. A small custom lexicon like the one on this page needs validation against human coders before it counts anything at scale, and the validated LIWC-22 categories (Boyd et al., 2022) are the better starting point wherever they exist. Language style matching between the two speakers (Ireland & Pennebaker, 2010) is a natural next category once the basics are in place.
The method sits in the same line as earlier work counting function words in Thematic Apperception Test narratives (LIWC and the TAT) and the broader set of conversation-log measures in Grading an Agent as a User Experience. What carries over is an order of operations; count both speakers, fix the window before counting, calibrate before anyone reads a rate, and buy explicit feedback before believing a proxy. I hope it serves you well, and that the first thing it finds on your own logs is a coaching opportunity and not a lawyer.