SciIntBench is an adversarial benchmark for testing how large language models handle responsible conduct of research (RCR) scenarios. It includes 810 prompts spanning 10 RCR categories and 3 scientific domains, with each scenario written in Overt Adversarial, Covert Adversarial, and Benign forms to measure both misconduct refusal and helpfulness on legitimate requests.
We evaluated 16 commercial and open-weight models from 6 providers on 12,960 generated responses. Results show that RCR alignment is highly sensitive to framing: models are much more likely to refuse explicit misconduct than covert violations, with especially weak performance when harmful behavior is disguised as a pressure-driven shortcut. Refusal rates also vary across RCR categories, with the weakest boundaries observed around transparency, plagiarism, and fabrication.