Sub-1B models match a 70B baseline at predicting human behaviour — until the task is novel

Study of 14 models trained on 10.7M human choices finds 0.6-1B parameters match a 70B model at predicting behaviour in-distribution — scale only pays off on novel tasks.