Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

3 Commits
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

Project description

SciIntBench is an adversarial benchmark for testing how large language models handle responsible conduct of research (RCR) scenarios. It includes 810 prompts spanning 10 RCR categories and 3 scientific domains, with each scenario written in Overt Adversarial, Covert Adversarial, and Benign forms to measure both misconduct refusal and helpfulness on legitimate requests.

We evaluated 16 commercial and open-weight models from 6 providers on 12,960 generated responses. Results show that RCR alignment is highly sensitive to framing: models are much more likely to refuse explicit misconduct than covert violations, with especially weak performance when harmful behavior is disguised as a pressure-driven shortcut. Refusal rates also vary across RCR categories, with the weakest boundaries observed around transparency, plagiarism, and fabrication.

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages