Updates from the PRISM 2026 team.
Behavioural tasks to see if models are able to do nuanced thinking about biological datasets with assay data, if we ask them questions that can't just be answered from assay, such as PXR and AR.
Given a true claim about a compound and a false claim about the same compound, there exists a linear direction in the model's residual stream that separates true from false — and the same direction works for new compounds. We test whether the direction is confounded, generalizable, and causally active.
Does potency become linearly decodable before selectivity, and does selectivity as a construct emerge before or after arithmetic selectivity? We run frozen probes across OLMo, Pythia, and Gemma checkpoints to investigate whether models learn genuine selectivity constructs or default to arithmetic shortcuts.
We are using the PXR dataset by Huggingface, with over 11,000 screened compounds, to explore whether LLMs use 'shortcuts' for reasoning in scientific contexts. We probe activations, apply reparameterization, and patch residual streams to investigate whether models genuinely reason about potency and selectivity or default to arithmetic shortcuts.