A direction in a network's activation space that reliably corresponds to one interpretable concept, rather than one neuron.
Continue to AI University →