What formalisms are good for defining good?
Working proposition: We may want a formalism for “good” that can be used in AI alignment. Question: good for whom, relative to what, and on what timescale?
Working proposition: “Good” may not be adequately represented by a single number. Question: what structure would be better, and for which decisions?
Working proposition: “Good” may have non-arbitrary structure. Question: what kind of structure, and which invariants survive across attempts to formalize it?
Research prompt: which transformations should preserve a judgment of good? Question: could invariances supplement content-based definitions? The Noether comparison is an analogy, not a proposed theorem.
Counter-prompt: perhaps judgments of good change with context in structured ways. Question: is covariance a useful mathematical model, or only a metaphor?
Extra question: what, if anything, in human cognition and experience would be interesting or valuable to a much more capable system? What evidence could distinguish answers from projection?