“Human-centred.” “Autonomy.” “Accuracy.” These words appear in almost every AI ethics framework, every responsible innovation charter, every design methodology worth its name. They signal a commitment to the person over the system, to dignity over efficiency, to truth over convenience. In the landscape of AI development, they function as a kind of ethical common ground, something everyone agrees on before the harder conversations begin.
The harder conversations, it turns out, begin the moment you ask what these words actually mean.
Working within COMFORTage, a Horizon Europe project developing AI-based tools for dementia and frailty prevention, we encounter this question not as an abstract provocation but as a practical one. When we sit with the people who will design, deploy, use, and live with these technologies, we bring with us values from ethics and trustworthiness frameworks as our toolkit. But soon enough, as the co-design discussions start, the first thing we notice is that ethical values do not converge into shared definitions. For a social scientist, that divergence is the data. Values do not fail when they do not converge; instead, they reveal something: distinct orientations, each coherent, each grounded, each carrying a different implicit answer to what technology should ultimately do.
Take accuracy. It relates to one of the most powerful arguments made on behalf of AI in clinical settings is that a well-designed system can overcome the limitations of individual human judgement. AI systems do not get tired, do not carry unconscious preferences, do not vary from patient to patient. They produce results that are accurate, they adhere to a ground truth, a measurable benchmark against which the system’s outputs can be validated. This is objectivity operationalised – and it is a genuine and important promise.
In a recent COMFORTage project meeting, a developer and a physician were discussing exactly this. For the developer, ground truth was the labelled dataset, the target variable that defines what a correct prediction looks like. The model’s accuracy is measured against it. For the physician, the phrase landed differently. “Ground truth” he said, “is the wellbeing of the patient in front of me. That is what we should be measuring. That is what accuracy ought to mean”.
The exchange is not a misunderstanding to be cleared up with better communication. It is a substantive disagreement about what the system is for. A model can be highly accurate against its target variable and still be measuring the wrong thing, or measuring a proxy for wellbeing that reflects some patients’ experiences well and others’ poorly. The bias the technology was meant to eliminate may simply have moved upstream, into the choice of what to measure and whose experience shaped the training data. Ground truth, it turns out, is not found, it is chosen. And that choice carries values, about what a “correct outcome” looks like, what “health” means, and whose reality the system is built to reflect.
But it doesn’t stop there. Even once a ground truth is agreed upon, the data used to measure it is not neutral. Research from within the project points to a further complication: the quality of data produced by users (in this case older adults, family members, caregivers) is directly linked to their familiarity with digital tools. In the current state of many AI-assisted health systems, a significant portion of the older adults these systems are designed to serve are not yet able to produce data of the quality the model requires to function accurately. The system’s accuracy, in other words, becomes partly a function of digital literacy. It performs well for users who already know how to use it, and less well, or not at all, for those who most need it.
This means the bias the technology was meant to eliminate may have moved upstream twice over: first into the choice of what to measure, then into the unequal capacity to measure it. Ground truth now looks produced, unevenly, by people with unequal access to the tools that produce it.
Autonomy fractures differently. The design specification for a predictive health tool promises that “the system supports patient autonomy through personalised recommendations”. Reasonable enough, on paper. But an older adult in a consultation says: “I don’t want to be told what to eat by a machine. I want my daughter to sit with me at lunch.” A neurologist explains that autonomy for her patients means slow explanation, checked understanding, family involvement. A caregiver of a person with dementia says: “He can’t choose for himself anymore: autonomy means I know him well enough to choose as he would have.” A policymaker frames it as individual responsibility for managing health autonomously, through digital tools. A citizen reframes it entirely: “Why does autonomy mean I have to change my behaviour, instead of society organising care and environment so I don’t bear the whole burden alone?”.
Five people, five definitions, each coherent, each pointing toward a different system design.
Human-centredness carries a different kind of difficulty. It presents as self-evident (of course the human should be at the centre) which makes it harder to contest and easier to use as a placeholder. A UX designer understands it as a methodological commitment: test with users, iterate, remove friction. A geriatrician understands it as relational: the technology should support the care relationship, not interrupt it nor overburden it. A family caregiver hears something else: “Human-centred would mean the system understands that I am also a human at the centre in this”. An older adult shown a prototype says quietly: “You’ve made it very easy to use. But I didn’t ask for this. Human-centred would mean someone asked me first”. A public health researcher points out that centring “the human” without asking “which humans” risks becoming a rhetorical blank cheque. The phrase sounds principled but its very openness is what makes it useful to those who prefer not to be too specific. When no one defines which humans are at the centre, the definition tends to fill itself in quietly with whoever is already visible, already resourced, already producing data that fits the existing frameworks. The risk is convenience over value.
What connects these three cases is not simply that values are interpreted differently by different people. That is unsurprising. The more important observation is where the differences come from. They are not the result of stakeholders’ confusion or insufficient ethics training. They emerge from genuinely different positions within a complex system: different professional traditions, different relationships to risk, different experiences of what it means to need care or to provide it.
This points toward a different understanding of what ethics in technology development actually requires. Philosophical analysis can clarify concepts, identify contradictions, and articulate principles. What it cannot do, on its own, is tell you how autonomy is lived by a caregiver making decisions for someone who can no longer make them, or what ground truth means to a clinician whose benchmark is a person rather than a variable. For that, you need something closer to fieldwork, the patient methods of social research, the willingness to sit with people in their actual contexts, and the discipline to treat a divergence in meaning as evidence rather than noise.
This is not a critique of philosophical rigour. It is a critique of all who think that integrating one social scientist or one philosopher in the project can solve it all. It is an argument about what ethics needs alongside it: time and effort for empirical inquiry that takes seriously how values are actually held and lived, rather than assuming they can be defined once and applied universally.
In the COMFORTage project, this is implemented through our social acceptability assessment. As a start, we treat this social angle of our work not as a compliance requirement but as an open research question. When we engage those who will live with the consequences of an innovation, we should include those who are ambivalent, resistant, or simply absent from the design table — what looked settled on paper reveals itself as uncertain. What seems obvious may mean something entirely different in practice. Responsible innovation, in this view, is not a checklist to be completed before deployment. It is a continuous process of listening, revising, and remaining open to the possibility that the values we started with need to be renegotiated in light of what we find.
Ethics that does not go looking for those differences is not neutral, it simply inherits the assumptions of whoever was already in the room.