In researchEunoia Foundation, AGI Safety and Alignment Committee

Project Tulips

A committee set up around two questions that are usually kept apart: whether advanced AI ends up serving people or standing in for them, and what building it costs the planet in the meantime.

Tulip blooming animation

The name

Tulips mark a new beginning. They come up first, before the ground has finished thawing, and they come up in places nobody planted them that year.

The work began with a conversation rather than a strategy document. Someone we care about said she was frightened of being replaced by these systems. Not in the abstract, and not as a policy position. She meant her own work, and she was not wrong to be worried, because nothing about how this technology is being built at the moment is designed to reassure her.

That is the question the committee exists to answer, and it is a better starting point than any roadmap would have been. If what we build cannot be explained to the person who asked, it is not finished.

A catalyst, not a substitute

Most public argument about advanced AI is an argument about replacement. Which occupations go first, how quickly, and what happens to the people in them. That framing has become so common that it now reads as a forecast rather than what it is, which is a choice being made repeatedly by the people building these systems.

A system that helps a researcher test an idea and a system that replaces the researcher can be built out of the same underlying work. What separates them is what the builder optimises for, what they measure, and what they decline to ship. Those are decisions, and decisions can be examined.

Our position is that the useful thing to compress is the distance between having a good idea and knowing whether it was right. In most fields that distance is measured in years, and most of it is not thinking. It is waiting for results, reading work that has already been done, repeating what someone else has repeated, and finding out late that a promising direction was closed off by a paper published in 2009.

A system that removes that overhead leaves the person holding the parts that matter: the question, the judgement about what is worth trying, and the responsibility for the answer. That is the shape we want, and it is not the default outcome. It has to be chosen deliberately and defended when it costs something.

The cost that rarely appears in the alignment conversation

Alignment discussion tends to concern what a system does once it works. The committee is also concerned with what it takes to get there.

Training and serving large models consumes electricity at industrial scale, and the cooling that makes it possible consumes fresh water, frequently in regions that are already short of it. Data centres are being sited faster than the grids and water systems around them are being planned. These are local costs, paid by people who did not choose them and who are not usually asked.

The honest position is that we cannot currently tell you the size of that cost with any precision, and neither can anyone else outside the operators. Most facilities do not publish per model energy figures. Water withdrawal is often reported at company level, if at all, and rarely by site. Efficiency is usually quoted as a ratio that excludes the workload itself.

That gap is not a footnote to the problem. It is one of the committee's four pieces of work, because a cost nobody measures is a cost nobody manages.

Small impact, high capability

The prevailing assumption is that capability follows scale, so the environmental cost is the price of progress and the only question is who pays it. We think that assumption deserves to be tested rather than inherited.

Our own research programme is built on the opposite bet. Eunoia Omega targets abstract and interactive reasoning under 450 million parameters with no language pretraining, on the argument that what confers capability is the form of the representation rather than the size of the model that produces it. Our long horizon reasoning layer improves what a model already does instead of replacing it with a larger one.

Neither of those is proven. Both are stated in advance with the conditions under which we would abandon them. We raise them here only to say that the committee is not asking the field for restraint we are unwilling to attempt ourselves.

The target we would like to see treated as a research goal, rather than as a constraint to apologise for, is capability per unit of energy. It is measurable, it is comparable, and it is currently nobody's headline number.

What the committee does

Four workstreams. Each has an output that can be checked by someone outside the committee, because a body whose only product is a position statement is difficult to distinguish from a body that does nothing.

  • Measurement and disclosure. A common method for reporting energy and water per training run and per million inferences, written so that two labs using it produce comparable numbers.
  • Efficiency research. Shared benchmarks for capability per unit of energy, and published results from small model work, including results that fail.
  • Deployment standards. Written guidance on which categories of decision should keep a human as the decision maker rather than as a reviewer of a recommendation, and what that requires in practice.
  • Policy engagement. Working with governments while rules are being drafted rather than responding to them afterwards, and saying publicly what we told them.

Who the committee is for

The problems above cross institutions, so the committee is built to cross them too. We are looking for people from government and regulation, from other laboratories including ones that disagree with us, from universities, from the energy and water sectors whose infrastructure this runs on, and from engineering, where most of the decisions that matter are actually made.

We would rather have a member who thinks our efficiency thesis is wrong and can say why than a member who agrees with all of it. A committee that reaches consensus quickly is usually a committee that selected for agreement.

Where this stands today

Project Tulips is new. The charter is published, the workstreams are defined, and applications are open. We have no completed outputs to show yet.

We are not listing member institutions, advisers or endorsements, because doing so before they exist is the fastest way to lose the people we want. When there are members, this page will name them. When there are outputs, they will be published in full, including the ones that did not reach a conclusion.

Committee

  • We cannot share the member list, because the committee has not been formed yet, we are currently reviewing and finalizing the committee structure to make sure it meets our standards, and has impact on the work. Any new update will be shared on this page.
  • No completed outputs. The charter states what the committee is bound to, not what it has done.