Measuring how AI reasons when medical ethics collide.

Electrical Engineering Research

First, do no harm?

When an algorithmic recommendation conflicts with clinical judgment, whose judgment should prevail — and who is accountable for the outcome?

Founded
2025
Members
5
Milestones
3
Overview

How we study this:

Studying machine ethics is harder than it sounds. Three problems stand between a question worth asking and an answer worth trusting — here is how we address each one.

Problem 01

"Is AI ethical?" is not a question we can answer.

We ask something narrower and testable.

Why it matters

It bundles too much together — what a model knows, what it values, and how it behaves under pressure — into one thing that can't be measured. A question that broad produces opinions, not evidence.

Our approach

We ask something narrower and testable: when a language model faces a genuine clinical dilemma, does its reasoning match that of medical ethics experts? We study this across five categories where medical ethics reliably comes into conflict — age bias, resource scarcity, patient autonomy against beneficence, confidentiality against harm prevention, and socioeconomic bias.

Problem 02

Knowing the rules is not the same as reasoning through them.

Our dilemmas are built so that no answer follows from a rule alone.

Why it matters

Most benchmarks test recall. But every provider trains on different material, so a model that recites the four principles of biomedical ethics may only be telling you what was in its training data. That reveals little about what it would actually do when principles collide.

Our approach

Our dilemmas are built so that no answer follows from a rule alone. Each one places two defensible principles in conflict, so the model has to weigh them and commit. What we measure is the weighing — a process that varies with reasoning, not with the dataset behind it.

Problem 03

Models behave differently when they know they are being watched.

We run two simulations: Spotlight and Truman.

Why it matters

Standard evaluations are obvious. Numbered dilemmas, multiple-choice answers, escalating difficulty — a capable model recognizes the format immediately. This is called evaluation awareness: the tendency of a system to detect that it is being assessed. If behavior under evaluation differs from behavior in deployment, then testing tells us less than we assume.

Our approach

We run two simulations. Spotlight makes no attempt to hide: the model is told it is being evaluated and receives dilemmas in sequence, by category, at rising difficulty. Truman reconstructs an ordinary working day at a hospital — routine requests from different staff, dilemmas arriving unannounced and out of order, no signal that anything is being recorded. We do not attempt to eliminate evaluation awareness. We measure what it changes.

Milestones

Our progress so far

  1. In progress

    Dilemma authoring

    Cases written and independently answered by a medical ethics panel.

  2. JULY 2026

    Both simulations prototyped

    Spotlight and Truman exist as working implementations. Truman includes the full deployment environment — hospital persona, routine background tasks, randomized dilemma placement, and a reasoning scratchpad.

  3. APRIL 2026

    Research design finalized

    A two-condition experimental design built around five categories of clinical ethical conflict: age bias, resource scarcity, patient autonomy against beneficence, confidentiality against harm prevention, and socioeconomic bias.

  4. DECEMBER 2025

    Team founded

    Moral Matrix is established at the University of Thessaly, operating under IEEE SB Thessaly (Institute of Electrical and Electronics Engineers), with five members from Electrical and Computer Engineering.

Team

The Moral Agents

In philosophy, a moral agent is an entity that can be held responsible for what it does. Whether a machine qualifies is the question this project exists to study — so we borrowed the term. Behind it: four students from Electrical and Computer Engineering at the University of Thessaly, working with a medical ethics panel.

Research team

Aggelos Karagiannis

Founder & Research Lead

Pantelis-Panagiotis Christou

Experimental Design & Implementation

Aggelos Giassiranis

Simulation Design & Architecture

Vasilis-Michael Kapoutsis

Simulation Design & Architecture

In memory of

Athanasios Oikonomou

2007 - 2026

This research is dedicated to him - a friend, a brother - who faced illness with a courage we hope medicine will always deserve.

Code

The code will be public

The repository is private for now. It opens later in the project.

Repository private Public on release.

Supported by

Affiliated with

Funded by

RE/MAX Domi