Parafoil · Methodology
How we read a conversation
If you worked with a human coach, you’d ask where they trained and what they believe. Fair question. Here’s our answer.
Parafoil gives you private feedback on how you lead, from meetings you choose to analyze. That feedback is only useful if you can judge it. And you can’t judge it without knowing where it comes from.
So: where the science comes from, how we test whether it holds, and where the method is most likely to be wrong.
Parafoil has analyzed thousands of real working conversations across organizations. Everything below is built on those. None of it comes from scripted sessions, role play or simulated material, which matters more than it sounds: a method tuned on performances of meetings learns what a performance looks like.
Measuring leadership is harder than it looks
You can’t measure leadership directly, the way you measure a temperature. You have to break it into things you can observe, then show those things add up to something real. That step is where most measurement of human behavior goes wrong.
It goes wrong in two directions. A system can measure something very consistently that turns out to have little to do with leadership. Or it can name something real and measure it so erratically that the number is noise. Everything below is built against both.
There’s a third reason for care. Kluger and DeNisi’s 1996 meta-analysis of feedback research, covering several hundred studies, found roughly a third of feedback interventions left performance worse than no feedback at all. Feedback is an active intervention with a real failure rate. Vague feedback is where most of the damage happens. That result shapes this product more than any other.
Eight layers, and none of them work alone
No single signal makes a judgment. Each layer answers a different question, and they check each other. A reading one layer produces and the others don’t support doesn’t reach you.
The layers combine conjunctively. For anything that reaches you as a claim about a specific moment, the supporting layers have to agree. Where they disagree the claim is suppressed instead of averaged, and a strong reading on one layer can’t carry a weak one on another.
What was said, and how
Psycholinguistic reasoning
Established scienceEnglish has a few hundred function words and around a hundred thousand content words. Those few hundred account for more than half of everything we say: pronouns, articles, prepositions, auxiliaries, negations, quantifiers.
Decades of published work, much of it from James Pennebaker’s group at the University of Texas, shows these track attention, relative status, certainty and psychological distance. How often you say “I” moves with self-focus and with lower status in the room. Pronouns shift around topics people find difficult. Article and preposition density rises when someone thinks concretely and falls when they abstract away. Tentativeness markers cluster where commitment is weakest. These effects are documented across many languages, and the word-count instruments behind them have been built well beyond English.
This layer is arithmetic. Words are assigned to categories from published dictionaries and counted, so any figure it produces can be traced to the exact words that produced it. It is the one part of the system that is not a black box in any sense, and we put it first for that reason.
Semantic reasoning
Established scienceWhat the words meant. Whether a question sought an answer or performed agreement. Whether a commitment was made or gestured at. Whether an objection got answered or absorbed.
Language models do this well. On their own they also produce confident readings of things that didn’t happen. This layer proposes. It doesn’t decide.
Vocal intonation
Established scienceTone carries what a transcript throws away. Pitch, pace, pause length and emphasis separate a question from a challenge built of identical words. Research on vocal affect is long-standing and well evidenced. A transcript on its own is an incomplete record of a conversation.
What counts as leading well
Behavioral science models
Established sciencePeer-reviewed models of how people behave in groups. Psychological safety, and whether people will speak up, disagree or admit a mistake in front of you. Self-determination theory, and the finding that autonomy, competence and relatedness sustain motivation where pressure doesn’t. Goal-setting research, and the repeatedly replicated result that specific goals beat being told to do your best.
We take these as given and cite them internally, so any claim the product makes traces back to a source.
Management literature
Established scienceThe applied work on what good management looks like day to day. Delegation. One-to-one practice. How change survives its trip through a management layer. What happens to teams under ambiguity.
This evidence base is messier than the experimental literature, with weaker controls and more case study. We treat it that way. It gives us well-evidenced patterns. It doesn’t give us laws.
Behavioral primitives
Our own workThe bridge from that research to something countable. Small observable units of the kind managing is actually made of: how a decision gets made and who ends up holding it, whether feedback lands specifically enough to act on, where attention goes.
Each one is scored on a short ordinal scale with an explicit “can’t tell” option, so a rater who cannot tell never has to guess. Every primitive is defined in writing, with worked examples at each point on the scale, and the definition is frozen before anything gets scored against it. The frameworks you may recognize are built from these units, so a framework can change without disturbing the measurement underneath.
What keeps it honest
Expert human rater panel
Our own workExperienced managers, recruited for range across function, sector, region, gender and ethnicity, and paid for their time. They rate real anonymized moments against the primitives, blind to each other and with the moments presented in a randomized order.
Some items come round a second time without being marked. That gives us test-retest agreement for each rater, so we know whether someone agrees with themselves before we ask whether the panel agrees with each other. Raters also mark the span of text their judgment rests on, which lets us check that two people who gave the same score gave it for the same reason.
Their judgment is what the automated scoring gets tested against. This is the part of the method that’s ours, and the part that took longest to build.
What makes it yours
Longitudinal baseline
Our own workParafoil builds a picture of how you lead over time and measures you against it. This matters more than it sounds. Comparing you to other people carries every stable difference between you and them. Comparing you to your own history removes all of it.
That comparison has its own problems and we read it accordingly. A quarter with a reorganization in it looks different for reasons that aren’t you. Knowing you’re measured changes behavior on its own. Unusually high and low readings tend to drift back toward your average whatever you do. So movement is read over long windows, and a single step up or down is not treated as a result.
The same picture shapes how advice reaches you. Recommendations are phrased in the vocabulary and register you actually use, because a suggested sentence you would never say is no use to anyone.
Orchestration
Our own workWorking out which frameworks matter for this conversation, with this person, at this point in the relationship, and holding the right history while it does. A correct observation about the wrong thing helps nobody. This layer is the difference between analysis and advice.
The small words do the work
The first layer is the least intuitive of the eight, so here’s one sentence with the content words faded and the function words picked out.
I think we should probably hold the release until you have had a proper look at it
12 of 17 words carry no topic at all. That is 71% of what you just wrote.
Underlined function wordsFaded content words
The topic disappears. The structure stays. Who’s the subject of the sentence. Whether a decision is being announced or offered. Where the hedging sits, and who it’s aimed at. One sentence proves nothing. Several hundred turns hold steady enough to say something about how you lead.
Good according to whom?
Any system like this risks encoding somebody’s private opinion about how to be a person. Three answers.
Claims trace to published research. Where the product says a behavior matters, that comes from the behavioral and management literature and is cited internally to a source. Our taste doesn’t enter it.
You’re compared to your own organization, and to your own history. Built from how leadership actually works where you work, and from how you led last month. Being more direct than average means nothing until you know what average looks like in your context.
You set the target. At onboarding you nominate what you want to work on, and guidance weights toward the leader you said you wanted to be. That makes this closer to instrumentation than assessment.
The weighting applies to what gets surfaced, and never to the measurement underneath. Two managers behaving identically are scored identically. They may be shown different things about it, which is what keeps organization-level views comparable across people who set different goals.
How we know it isn’t making things up
The reasonable suspicion about any language system is that it writes plausible prose with nothing behind it. Three things stand against that.
Everything is scored against a human reference. The primitives aren’t self-assessed. They’re measured against the expert panel’s judgments on the same material. People can show the system is wrong. It can’t just mark its own work.
It abstains. When audio is poor, speakers can’t be separated reliably, or your own words are badly captured, you get no score. Returning nothing beats returning something confident and wrong. Abstention is a designed output with its own criteria.
Every claim is anchored. When Parafoil says something happened, it shows you the moment and the words. You can check any observation against what was actually said, which makes the whole thing testable by you.
Nothing reaches you until it’s reliable enough
Agreement between expert raters is measurable, with a standard toolkit behind it: chance-corrected agreement statistics for categorical judgments, intraclass correlation for scaled ones, and confidence intervals built by resampling so no single estimate gets read on its own. It’s the same apparatus used to validate clinical and psychological instruments.
When experienced managers independently reach the same read on the same moment, a behavior is defined well enough to score. When they scatter, the definition is at fault and the raters aren’t. The honest response is to fix the definition and hold the number back.
So each behavior sits on a ladder, and where it sits decides what the product may do with it:
- LowNot used at all. The definition needs work before the measurement means anything.
- ModerateAggregate patterns only. Dependable across many observations, unreliable on any one.
- GoodValidated. Used in analysis and in trends over time.
- High — the highest barIndividual-facing. Only here can a behavior drive feedback about one specific person.
The bar for saying something about you is the highest one. A behavior that hasn’t cleared it can still inform a pattern across hundreds of meetings, and still not tell you what you did in one conversation.
Small panels give wide confidence intervals, sometimes wide enough that a behavior sits across two rungs. When that happens the lower rung wins. The fix is more raters and better definitions, not a promotion.
How a behavior earns its place
Every primitive starts in the literature. It gets defined in writing, reviewed against the published research it came from, and frozen before a single conversation is scored against it. Then the panel tests whether experienced managers reading the same moment reach the same judgment.
A behavior has to survive both to reach the product, and it keeps having to. Definitions are re-tested as the panel grows, and a behavior that stops holding up goes back down the ladder.
The rule we hold ourselves to
The human ratings are an evaluation set and an entitlement gate. We never tune the scoring to agree with the panel more closely. A system trained against its own test gets a better score and becomes a worse product. The panel is there to tell us when we’re wrong.
Bias is the risk we watch hardest
A system that reads how people talk will reward some ways of talking. That’s a real exposure and we’d rather set it out than wait to be asked.
The risk is clearest for three groups: people speaking a second language, people from cultures where directness or deference work differently, and quiet people who contribute plenty and talk little.
What we do about it
- The panel is built for range. Raters differ by function, sector, region, gender and ethnicity as well as by management background. A reference standard inherits the blind spots of whoever sets it, which makes the composition of the panel part of the measurement.
- Comparison stays local. Measuring against your own organization and your own history removes a large class of cultural penalty that a global ideal creates.
- We score the outcome and stay agnostic about the route. Whether ownership of a decision moved is the measurement. Very different styles can achieve it and should read the same.
- Speaking time is context. It describes the meeting. It earns no credit.
- Coverage is monitored. We track how often the system declines to score a conversation, per person. A high rate is treated as a fault in capture quality and gets fixed, so nobody quietly receives thinner feedback than their peers.
The two we watch hardest
The principle travels between languages. The calibration is built per language, because languages that routinely drop subject pronouns, or carry status in honorifics, need their own mapping before the same inferences hold.
Second-language speakers are the case we watch most closely. Someone working in a shared company language can show patterns that belong to the switch and not to their leadership, so those readings are held to a higher bar before anything is said about them.
We monitor for systematic differences in scoring across groups, and we would rather hear from you than find it ourselves. If a read looks like it reflects how someone speaks and not how they lead, tell us. That’s a request, and we mean it.
Where Parafoil gets it wrong
The most useful thing we can give a skeptical reader is the shape of the errors.
It only sees the room
Parafoil reads what happened in the meeting. It doesn’t know you settled the question over Slack that morning, or that the person you were short with had asked you to be direct. When a read feels wrong, that’s usually why. You’re not obliged to accept it.
One meeting is weak evidence
Any single conversation is a small sample of an unusual day. The signal is in repetition: the same tendency across many meetings, with different people, over weeks.
It reads behavior, and stops there
It can’t see why you did something, and it makes no judgment about the kind of manager you are. Where the product names a style, it’s describing a pattern of behavior over time. That’s a much more modest claim than it sounds.
Most leadership feedback is someone’s memory
Almost everything a manager learns about how they lead arrives as a survey, a review cycle, or a colleague’s summary months after the fact. Each one is a recollection of what happened, collected once and compressed into a form.
The limits of that are well documented. People recall the recent and the dramatic and lose the ordinary middle, which is where most managing actually happens. What a colleague is willing to write down about you differs from what they noticed. A substantial share of the variance in multi-rater feedback reflects the person giving the rating as much as the person being rated. And it arrives once a year, about twelve months of behavior.
This is also where the feedback research bites. The interventions that leave performance worse are disproportionately the vague ones, and “be more strategic,” landing nine months after the meeting that prompted it, is close to the limit of vague.
Parafoil reads the conversation while it is still the conversation, and it reads every one you choose to include. The ordinary middle is the part it sees most of.
On coaching
None of that is an argument against a coach. A good coach does things Parafoil does not: they know your situation, they push back in the moment, and they hold you to what you said last time. Coaching and this are two methods pointed at the same goal, and they work well together.
What a panel can do that any single observer cannot is average. One person reads you through their own training, their own experience and their own good days, which is what being a person means. A reference built from many experienced managers spreads that across all of them.
What we’ve chosen not to do
All three of these are available to us. We’ve left them alone on purpose.
We don’t assess anyone but you
Parafoil captures whole conversations, so other people’s words are in the material. They aren’t the subject of it. Nobody else in your meetings gets a score, a profile, or a read of their own, and no view anywhere in the product produces one. The measurement is of how you lead. Everyone else in the room is context for that.
We don’t read faces
Video analysis of expression and posture is commercially available and some products use it. The evidence underneath it is weaker than the enthusiasm around it. Barrett and colleagues’ 2019 review of the field concluded that emotional state cannot be reliably inferred from facial configuration alone, and a measurement you can’t defend is worse than no measurement. We’d rather wait for science we can stand behind.
We don’t read your email or your chat
Your inbox and your messages would add real context. They would also quietly change what you agreed to. Parafoil reads the meetings you choose and nothing else, and that boundary fits in one sentence, which is most of why it can be trusted. We won’t add a sensor we can’t explain that simply.
Read the research yourself
Nothing on this page asks you to take our word for the science. The work below is published, and most of it is the standard reference in its own area. Pennebaker’s book is the general-reader way in; the rest is the academic literature the layers are built on.
Language and what function words carry
- Pennebaker, J. W. (2011). The Secret Life of Pronouns: What Our Words Say About Us. Bloomsbury Press.
- Tausczik, Y. R., & Pennebaker, J. W. (2010). The psychological meaning of words: LIWC and computerized text analysis methods. Journal of Language and Social Psychology, 29(1), 24–54.
Feedback, and why vague feedback is dangerous
- Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254–284.
What the behavioral literature says about leading
- Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383.
- Ryan, R. M., & Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1), 68–78.
- Locke, E. A., & Latham, G. P. (2002). Building a practically useful theory of goal setting and task motivation: A 35-year odyssey. American Psychologist, 57(9), 705–717.
Measuring agreement between raters
- Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378–382.
- Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.
On reading emotion from faces
- Barrett, L. F., Adolphs, R., Marsella, S., Martinez, A. M., & Pollak, S. D. (2019). Emotional expressions reconsidered: Challenges to inferring emotion from human facial movements. Psychological Science in the Public Interest, 20(1), 1–68.
Each title links to a literature search for that work, so you land on the paper itself and on everything that has cited it since.
More questions? Email humans@parafoil.co