Subject-Dependent Does Not Mean Arbitrary

Sensory experience depends on the perceiver, but that does not make judgment arbitrary. This essay examines what we mean by “subjective” in coffee—and how shared references, language, training, and comparison can make sensory judgments more reliable, communicable and intersubjectively meaningful.

Share
Subject-Dependent Does Not Mean Arbitrary

How private sensory experience becomes a professional claim

Two experienced tasters can describe the same coffee in similar terms and still disagree about its quality.

That possibility takes us a little further than the first Coffee Language essay.

There, the problem was description: how an experience becomes peach, apricot, or stone fruit, and why different words do not necessarily prove different perceptions. But once description becomes judgment, the question changes. If flavor depends partly on the person tasting—on attention, memory, learning, expectation, and context—what kind of claim can a sensory judgment make? If trained assessors disagree, is one judgment more defensible than another? And when we call one coffee better than another, are we discovering something in the coffee or applying a value to it?

The familiar answer arrives quickly:

Taste is subjective.

Sometimes that phrase expresses useful humility. No taster occupies a view from nowhere. We bring bodies, histories, habits, expectations, and learned categories to the cup.

But subjective can also become a way of ending the inquiry. If an experience belongs to a subject, the reasoning goes, then it is personal; if it is personal, nobody else can question it; and if nobody can question it, one response is ultimately as valid as another.

That conclusion does not follow.

Sensory experience is subject-dependent: it occurs for someone, through a particular perceptual system and under particular conditions. Arbitrariness is a different matter. A judgment becomes arbitrary when it is insufficiently constrained by the coffee, the task, the method, the criteria, or the evidence available to support it.

The presence of a subject does not remove those constraints.

Where flavor happens

A coffee can exist without being tasted. Its flavor as experienced cannot.

Long before evaluation begins, the coffee brings a material history to the encounter. Genetics, growing environment, agronomy, maturity at harvest, processing, drying, and storage have already shaped the green coffee. Roasting changes it again, and grinding, water, concentration, temperature, and preparation determine what finally reaches the drinker.

Liberman’s (2022) fieldwork repeatedly emphasizes how dynamic this material object is. Coffees intended to reproduce a stable commercial flavor still change across harvests because rainfall, temperature, disease, processing, storage, and other conditions do not remain fixed. Industry accounts have likewise long treated cultivation, cherry selection, processing, drying, and roasting as contributors to what eventually becomes coffee quality (Grieg-Gran, 2005).

The perceiver does not invent that material history. The beverage arrives with physical and chemical properties that constrain what can happen sensorially.

But those properties are not identical to the experience of flavor.

Flavor perception involves the integration of gustatory, olfactory, somatosensory, thermal, and chemesthetic information within a perceiving organism (Small & Prescott, 2005). What becomes noticeable can also vary with attention, adaptation, previous learning, expectation, and context. The same beverage may therefore yield somewhat different experiences when either the material conditions or the conditions of perception change.

There are philosophical positions that locate flavor properties more strongly in foods and beverages themselves. Smith (2012), for example, argues that flavors can be understood as configurations of properties in foods that our sensory systems track, while recognizing that experiences of those properties vary. We do not need to settle that ontological question here.

For sensory evaluation, a narrower distinction is enough: the chemistry of the coffee and the experience of the coffee are not interchangeable descriptions of the same thing.

To call an acidity sharp is not to claim that sharpness sits inside the beverage like caffeine or an organic acid. It describes how something about that beverage becomes perceptually organized for someone under particular conditions.

That judgment may be repeatable or unstable, well supported or careless. Other assessors may recognize it or reject it. They may even agree about the sensory character while understanding sharp somewhat differently.

Those are questions about the quality of the judgment. Simply saying that a subject was involved does not answer them.

What objectivity can mean

The objective-subjective opposition appears straightforward until we ask what objective is supposed to mean.

Sometimes it means existing independently of any perceiver. Sometimes it means relatively independent of personal preference. It may refer to a method whose procedures are public, constrained, and repeatable, or more loosely to something expressed through numbers or instruments.

Those meanings overlap, but they are not equivalent.

Consider a moisture measurement. People selected the property to be measured, defined the units, designed the instrument, established calibration procedures, and decided what level of uncertainty would be acceptable. Human involvement does not make the result arbitrary. Confidence comes from the way the measurement is structured: the target is specified, individual discretion is restricted, sources of error can be examined, and others can check the result under stated conditions (Maul et al., 2018).

This does not place an instrumental reading and a sensory judgment on the same epistemic footing. A calibrated instrument can achieve a degree of traceability and independence from individual perception that a taster cannot.

Sensory evaluation presents a different problem because the subject is part of the phenomenon being reported. Remove the taster and an instrument may continue to register pH, concentration, or temperature. Remove the taster and there is no experienced sweetness, jasmine, sharpness, or pleasure left to report.

The goal cannot therefore be to eliminate the subject. It is to understand what the subject contributes and to constrain the judgment enough that others can examine what it means and how much confidence it deserves.

Kenneth Liberman’s Tasting Coffee is particularly useful here because he treats objectivity not mainly as a philosophical abstraction but as something coffee professionals have to accomplish in practice. A roaster, importer, or buyer may need a flavor profile to remain sufficiently recognizable across years, origins, shipments, and people. The industry cannot pretend that tasters are neutral machines, but it also cannot operate as though every sensory response were beyond comparison.

For Liberman (2022), professional objectivity is maintained through repeated tasting, shared procedures, calibration, communication, and continual work to keep descriptors and scales usable across people. His fieldwork repeatedly shows that these systems do not maintain themselves. Tasters have to do that work. 

Numbers are part of it, but numeration does not remove the need for interpretation. A score becomes comparable across assessors only when they have sufficiently similar understandings of what the scale represents and how it should be used. Numerical systems can restrict some forms of discretion and support memory and comparison, but their professional meaning still depends on the practices surrounding them.

Different claims require different support

Part of the confusion around subjectivity comes from treating every human response as though it were making the same kind of claim.

Consider four statements:

I prefer cup A to cup B.

Cup A tastes more bitter than cup B.

Cup A is the better coffee.

Cup A contains 1.35 percent total dissolved solids.

These statements are not simply four points along a scale from subjective to objective. They are doing different things.

The first is a report of preference or hedonic response. If I sincerely prefer cup A, another person cannot disprove the existence of that preference by preferring B.

The second is a comparative sensory claim. The experience remains mine, but the judgment can be examined. Present the samples blind again. Repeat the comparison. Use reference solutions. Compare the ranking with those of other trained assessors. If the judgment repeatedly reverses when the coded samples return, there is less reason to treat the original statement as evidence of a stable sensory difference.

The third statement is evaluative. Calling A better introduces criteria. Better in what respect? For which use? According to whose standards? A quality judgment can be disciplined, but the grounds of that judgment need to be visible.

The fourth statement is instrumental. It depends on a device, sampling method, calibration, and operating conditions. Its uncertainty and evidential structure are different again.

Calling the first three subjective and the fourth objective obscures these differences. We need to know what kind of claim is being made and what evidence would be appropriate to it.

The distinction between perception and preference is especially important. Two tasters can agree quite closely on the intensity of acidity and still differ in how desirable they find it. They may agree that acidity is high while disagreeing over whether that intensity is appropriate, balanced, or excessive for the purpose at hand.

This is familiar at cupping tables. Sometimes two cuppers seem to disagree until they explain what produced their scores. The sensory observation may be similar while the weighting differs: one assessor may place more emphasis on clarity or flavor, another on acidity or mouthfeel. What first appears to be perceptual disagreement may actually sit elsewhere in the judgment.

Disagreement can therefore enter at several places: perception, description, scale use, preference, criteria, or purpose. Before trying to resolve it, we need to know where it occurred.

Trained subjectivity

Sensory science does not need to decide whether a person is subjective or objective in the abstract. It can ask more answerable questions.

Can the assessor discriminate among samples? When the same sample returns, does the response remain similar? Does one assessor use a scale more severely than another? Do panelists rank products in a similar order? Are their scores close enough for the decision being made?

Panel-performance methods examine these kinds of behavior through measures of within-assessor repeatability, between-assessor agreement, discrimination, scale use, severity, and assessor-by-product interactions (Bi, 2003; Latreille et al., 2006; Rossi, 2001).

None of those measures gives us direct access to another person’s private experience. They examine how a judgment behaves once it has been expressed.

This is where I find trained subjectivity useful as a working term.

I do not mean that training converts subjectivity into objectivity. The subject remains. What changes is the discipline around the response.

Training can teach a person where to direct attention, how to compare a sensation with a reference, how to use a scale, how to distinguish liking from description, and how to repeat a task under controlled conditions. It can also expose an assessor to situations in which an initial judgment fails and has to be reconsidered.

In teaching, I often see that the first change is not that someone suddenly tastes an entirely different coffee. A reference becomes clearer, scale use becomes more consistent, or the person becomes better at separating what they perceive from how much they like it. Over time, those changes can substantially improve the quality of the judgment.

The effects are task-specific. Chambers et al. (2004), for example, followed seven screened panelists evaluating tomato sauces after different amounts of training. Performance generally improved as training increased, with better discrimination and reduced variation between panelists, but the amount of training required was not the same for every attribute.

That is a useful limitation.

A trained taster may be highly repeatable on one sensory dimension and less so on another. Someone may use a scale consistently while remaining linguistically idiosyncratic. A credential tells us something about preparation for a task, but it cannot replace evidence of how an assessor actually performs in the situation at hand.

Liberman’s (2022) ethnographic observations point in the same direction. Professional objectivity is not permanently acquired once someone becomes an expert. Descriptors have to be restabilized, scales recalibrated, and practices checked against the coffees being tasted. He describes calibration as continual local work rather than a state that is achieved once and then possessed. 

Training, then, does not remove interpretation. It can make particular acts of attention, comparison, scale use, and judgment more disciplined.

What coffee evidence allows us to say

Research involving trained coffee assessors does not support a simple verdict.

Pereira et al. (2019) studied seven Q Graders evaluating coffees across different quality levels. Their results showed strong precision in some parts of the scale, particularly among coffees categorized as excellent and outstanding, while errors increased around an important classification boundary. The judgments were not random, but their stability was not uniform across the task.

That matters because boundaries have consequences. A relatively small scoring difference near a threshold can change a coffee’s classification and potentially affect how it is treated commercially.

Fioresi et al. (2023) examined concordance among Q Graders and found disagreements in attributes including flavor, acidity, body, and overall assessment, with some differences affecting final classification. Their analysis also highlights a distinction that is easy to overlook: assessors can be correlated without being concordant. Scores can move in the same direction across samples while remaining meaningfully separated in their absolute values.

Both studies predate the current CVA-based Q Grader program and should therefore be read as evidence about the protocols and tasks they studied, not as a direct evaluation of the current system (Specialty Coffee Association, 2025).

The term clean cup offers an even sharper example.

Giacalone et al. (2020) asked coffee professionals to evaluate clean cup on a line scale. Assessors could be relatively stable with themselves while agreement between assessors remained very low. The problem was therefore not simply random responding. Some participants were individually consistent while apparently relating to the construct differently from their peers.

A panel average can conceal this kind of disagreement. If panelists rely on different criteria or use different regions of a scale, the mean still provides a decision rule. It tells us much less about whether the judgments underneath it share the same meaning.

These studies involved particular panels, coffees, attributes, and methods. They should not be turned into a verdict on Q Graders, coffee professionals, or sensory evaluation generally.

What they do show is that reliability has to be demonstrated in relation to the task. An assessor can be repeatable on one attribute and less stable on another. A panel can discriminate samples while disagreeing about scale levels. A method can be sufficient for one decision and inadequate for another.

The useful question is not whether coffee judgment is “really subjective,” but where the judgment holds, under which conditions, and for what purpose.

From trained subjectivity to shared judgment

Individual discipline solves only part of the problem. Sensory work becomes professionally useful when judgments can move beyond the person who made them.

The moment a sensory experience becomes a professional claim, language enters the problem. A perception has to be named, scaled, scored, or otherwise expressed before another person can examine it. For sensory judgments to become comparable across people, it matters not only what tasters perceive, but also whether the systems through which they express those perceptions carry sufficiently shared meanings. This becomes especially visible during calibration. I have sat at tables where two people initially appeared to disagree and, after talking through the cup, it became clear that their descriptions were pointing to similar features of the coffee with different words. The opposite also happens: two assessors use the same descriptor and only later discover that they are referring to somewhat different sensations, intensities, or references. The word itself can create an appearance of agreement before the underlying meaning has been checked.

Liberman describes calibration as involving a shared intersubjectivity of both palate and language. Tasters have to coordinate references, descriptors, scales, and the sensory objects toward which those terms are directed. He uses intersubjective adequacy for the practical achievement through which tasting accounts become sufficiently coordinated to work beyond the individual (Liberman, 2022). 

There is a tension inside that practice.

Professional tasters often value independent, silent assessment because another person’s words or scores can influence the first judgment. Yet Liberman also observed tasters relying on collaborative tasting to notice qualities one person may have missed, test interpretations, and stabilize the use of descriptors. 

These activities serve different purposes. Independent tasting protects the initial observation from immediate social influence; later comparison can reveal how differently the same sample, descriptor, or scale has been understood.

I use structured intersubjective agreement for the support that can emerge when independently formed judgments are compared under shared conditions, procedures, references, and purposes.

This is not a claim that assessors have achieved identical private experiences, and agreement by itself does not prove that a group is correct. It means that the judgment is no longer supported only by the sincerity of the person making it. It can be compared with other observations and examined against a common task.

At the individual level, trained subjectivity asks whether a person has learned to make a more disciplined sensory judgment. At the collective level, structured intersubjective agreement asks whether independently formed judgments align sufficiently for the decision being made.

The fuller problem of calibration—how consensus is reached, how hierarchy influences a panel, and when agreement becomes conformity—deserves its own treatment. For the present argument, it is enough that subject-dependent judgments can become collectively examinable without becoming observer-independent facts.

Quality belongs to a relation

Any serious account of coffee quality has to begin with the coffee itself.

Coffee does not reach the evaluator as blank material. Genetics, environment, agronomy, cherry maturity, post-harvest processing, drying, storage, roasting, and preparation have already shaped its physical and chemical possibilities. An evaluator cannot simply imagine a floral aroma into a coffee with no relevant sensory basis for it, nor can preference undo the material consequences of severe defects or poor preparation.

Liberman’s (2022) account of coffee as a living and changing agricultural product is useful here. A company trying to reproduce a recognizable flavor may have to substitute origins, alter blends, or change purchasing decisions because harvest conditions have changed what particular coffees can provide. The need for reliable judgment becomes especially clear when the object itself is not static. 

Material constraint, however, is not yet the same thing as quality.

To move from this coffee has high acidity to this acidity is excellent introduces a criterion of merit. It asks not only what is perceived, but what that perception should count as.

An intense acidity may be valued in one coffee and judged inappropriate in another. A heavy body may serve one style or preparation and obscure another. Consistency may be a central quality for a commercial blend, while novelty and distinctiveness can carry greater value in a competition or auction.

I was reminded of this recently while judging the Taiwan Coffee Auction. At that level, jurors may agree quite readily that a coffee is exceptional and still disagree about how exceptional: whether it belongs around 90 points, for example, or several points higher. That difference does not necessarily mean that they perceived the coffee fundamentally differently. It may reflect differences in scale use, weighting, or in how particular qualities are valued at the upper end of evaluation.

One of Liberman’s (2022) more revealing observations is that professional coffee purveyors are not always searching for an abstractly “best” coffee. They may need the best coffee capable of reproducing a required flavor consistently, in sufficient volume, and for a particular customer or product. What counts as excellence is partly shaped by what the coffee is expected to do.

This points toward a relational account of quality.

I do not mean that quality is whatever anyone wants it to be. Coffee quality cannot be exhausted either by listing the material properties of the coffee or by asking whether an individual likes it. Quality judgments arise when those properties are experienced and evaluated according to criteria, purposes, and contexts that need to be stated rather than left implicit.

A professional quality judgment can therefore be disciplined without being universal. Preparation can be standardized, criteria can be made explicit, assessors can be trained and retested, and judgments can be examined for consistency and suitability.

But criteria still enter, and once they do, questions follow about who established them, what they are intended to accomplish, and whether they are appropriate to the decision being made.

Professional evaluation and consumer liking make the distinction especially clear. Giacalone et al. (2016) compared consumer responses to two coffees that had been valued very differently by professionals. Consumer preference did not simply reproduce the professional hierarchy. That does not invalidate professional quality assessment; it indicates that the professionals and the consumers were not necessarily answering the same question.

A coffee can perform well within a professional quality system without becoming everybody’s preferred cup.

Sensory quality, commercial suitability, market value, reputation, and personal liking interact with one another, but they are not interchangeable.

A fuller account of coffee quality will need to examine those relations in much greater depth. For now, the point is more limited: the coffee constrains what can reasonably be perceived and claimed; the perceiver makes flavor experience possible; and evaluation introduces criteria and purpose.

Beyond the binary

The objective-subjective divide gives us two convenient places to locate sensory truth: in the object or in the person. Coffee does not fit comfortably into that division.

Its material properties are real and consequential, but flavor as experienced requires a perceiver. The perceiver contributes attention, memory, learning, and expectation without having unlimited freedom to experience anything whatsoever. Training can make a judgment more disciplined without removing interpretation, and agreement can give a judgment wider support without turning it into a view from nowhere.

This is why subject-dependent is a more useful starting point than subjective. It tells us something important about where sensory experience occurs without deciding in advance how much confidence the resulting judgment deserves.

That confidence depends on how the judgment was produced: the conditions of tasting, repeated observation, the distinction between description and preference, the assessor’s performance in the relevant task, and, where the claim has to travel beyond one person, the extent to which independently formed judgments can be compared under shared conditions.

Quality complicates the picture because criteria and purpose enter as well. The coffee matters, but it does not contain a complete system of value inside itself. The evaluator matters, but preference alone does not account for professional quality judgment.

So rather than asking whether coffee quality is simply objective or subjective, I find another question more useful:

How was this claim about quality built, what constrains it, and what can it responsibly support?

That question does not resolve everything. It leaves open how standards acquire authority, how calibration should be organized, how language can clarify or conceal disagreement, and whose values become embedded when an industry decides what counts as quality.

Those are questions I want to return to.

For now, a sensory judgment can belong to a subject without belonging only to that subject. Training, repetition, comparison, and shared procedures can give it support beyond private experience.

It does not need to become absolute to become accountable.

Christopher Pearse Cranch, “Transparent Eyeball,” c. 1837–39, after Ralph Waldo Emerson’s Nature.

References

Bi, J. (2003). Agreement and reliability assessments for performance of sensory descriptive panel. Journal of Sensory Studies, 18(1), 61–76. https://doi.org/10.1111/j.1745-459X.2003.tb00373.x

Chambers, D. H., Allison, A.-M. A., & Chambers, E., IV. (2004). Training effects on performance of descriptive panelists. Journal of Sensory Studies, 19(6), 486–499. https://doi.org/10.1111/j.1745-459X.2004.082402.x

Fioresi, D. B., Ramos, A. C., Bertolazi, A. A., & Pereira, L. L. (2023). Adherence and concordance among Q-Graders in the sensory analysis of coffees. Journal of Sensory Studies, 38(2), e12805. https://doi.org/10.1111/joss.12805

Giacalone, D., Fosgaard, T. R., Steen, I., & Münchow, M. (2016). “Quality does not sell itself”: Divergence between “objective” product quality and preference for coffee in naïve consumers. British Food Journal, 118(10), 2462–2474. https://doi.org/10.1108/BFJ-03-2016-0127

Giacalone, D., Steen, I., Alstrup, J., & Münchow, M. (2020). Inter-rater reliability of “clean cup” scores by coffee experts. Journal of Sensory Studies, 35(5), e12596. https://doi.org/10.1111/joss.12596

Grieg-Gran, M. (2005). From bean to cup: How consumer choice impacts upon coffee producers and the environment. Consumers International & International Institute for Environment and Development.

Latreille, J., Mauger, E., Ambroisine, L., Tenenhaus, M., Vincent, M., Navarro, S., & Guinot, C. (2006). Measurement of the reliability of sensory panel performances. Food Quality and Preference, 17(5), 369–375. https://doi.org/10.1016/j.foodqual.2005.04.010

Liberman, K. (2022). Tasting coffee: An inquiry into objectivity. State University of New York Press.

Maul, A., Mari, L., Torres Irribarra, D., & Wilson, M. (2018). The quality of measurement results in terms of the structural features of the measurement process. Measurement, 116, 611–620. https://doi.org/10.1016/j.measurement.2017.08.046

Pereira, L. L., Guarçoni, R. C., Moreira, T. R., de Sousa, L. H. B. P., Cardoso, W. S., Moreli, A. P., da Silva, S. F., & Ten Caten, C. S. (2019). Very beyond subjectivity: The limit of accuracy of Q-Graders. Journal of Texture Studies, 50(2), 172–184. https://doi.org/10.1111/jtxs.12390

Rossi, F. (2001). Assessing sensory panelist performance using repeatability and reproducibility measures. Food Quality and Preference, 12(5–7), 467–479. https://doi.org/10.1016/S0950-3293(01)00038-6

Small, D. M., & Prescott, J. (2005). Odor/taste integration and the perception of flavor. Experimental Brain Research, 166(3–4), 345–357. https://doi.org/10.1007/s00221-005-2376-9

Smith, B. (2012). Perspective: Complexities of flavour. Nature, 486, S6. https://doi.org/10.1038/486S6a

Specialty Coffee Association. (2025, October 1). Q Grader courses now available across various regions and languages. https://sca.coffee/sca-news/2025/10/1/q-grader-courses-now-available-across-various-regions-and-languages-zakzc