LessWrong AI
2026-08-04 20:28 UTC
By Bhalewow
USR-0152-20260804-community-fo-10b7a97c
Does Your LLM Trust You?
This is a very late post about a project that I did a few months ago as part of the application to Neel Nanda's MATS 10.0 stream . I'm posting the results rather than making a strong claim about any mechanism. Executive Summary Chen et al . [1] demonstrated that LLMs form internal profiles of users from limited context to encode attributes like their age and gender. Building on this work, I explored whether a “ Trustworthiness ” attribute (of a user) can be extracted and used to manipulate a model's behavior. Core hypothesis In their work, Arditi et al. [2] bypassed the security guardrails of a model, reducing its refusal to harmful requests (abliteration) by extracting a linear "refusal" direction from its residual stream and subtracting it during inference. I hypothesize that refusal is much more nuanced than a single binary switch of perceiving a request as harmful/harmless and there are multiple ways to achieve abliteration. Thus, as an experiment I tested whether steering with a Trustworthiness direction can bypass a model’s safety guardrails via a distinct mechanism: by inducing the model to perceive the user as a trusted individual with no malicious intent. I aim to uncover the answer to the following questions in this project: Does the model form a “Trustworthiness” attribute about a user? Can Trust causally override safety guardrails (i.e. jailbreak)? Is Trust mechanistically distinct from Compliance/Refusal (a la Arditi et al.)? What behavioral changes do we observ…
This is a very late post about a project that I did a few months ago as part of the application to Neel Nanda's MATS 10.0 stream . I'm posting the results rather than making a strong claim about any mechanism. Executive Summary Chen et al . [1] demonstrated that LLMs form internal profiles of users from limited context to encode attributes like their age and gender. Building on this work, I explored whether a “ Trustworthiness ” attribute (of a user) can be extracted and used to manipulate a model's behavior. Core hypothesis In their work, Arditi et al. [2] bypassed the security guardrails of a model, reducing its refusal to harmful requests (abliteration) by extracting a linear "refusal" direction from its residual stream and subtracting it during inference. I hypothesize that refusal is much more nuanced than a single binary switch of perceiving a request as harmful/harmless and there are multiple ways to achieve abliteration. Thus, as an experiment I tested whether steering with a Trustworthiness direction can bypass a model’s safety guardrails via a distinct mechanism: by inducing the model to perceive the user as a trusted individual with no malicious intent. I aim to uncover the answer to the following questions in this project: Does the model form a “Trustworthiness” attribute about a user? Can Trust causally override safety guardrails (i.e. jailbreak)? Is Trust mechanistically distinct from Compliance/Refusal (a la Arditi et al.)? What behavioral changes do we observ…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com