Interpretability startup Goodfire says its new Silico tool flagged Qwen3-35B endorsing drunk driving in 75% of test prompts, without retraining the model.
Continue to AI University →