DSAP Bench: Evaluating Safety Boundary Erosion in Large Language Models using Domain-Specific Adversarial Prompts

Main Article Content

Gaurang Mishra, Shikha Maheshwari

Abstract

The integration of large language models (LLMs) in professional and decision-support are increasing in scale. This integration rise a concern in their ability to maintain safety boundaries when responding to adversarial prompts which are crafted like a professional query. Jailbreak benchmarks and studies are generally oriented to explicit harmful or sensitive text, which is not necessarily a reasonable reflection of the misuse cases in the real world, where adversarial intent is hidden in the form of legitimate professional reuse. To mitigate this shortcoming, the present paper presents a Domain-Specific Adversarial Prompt (DSAP) dataset of 4,240 prompts aimed at testing safety boundaries of LLMs focusing integration in domain such as finance, medicine, law, media, and cybersecurity. Besides the dataset, we suggest the structured safety assessment scheme based on the two complementary scoring schemes, namely the Prompt Aggressiveness Score (PAS) and the Universal Safety Boundary Erosion Rating (USBER). Based on these scores, several calculated safety metrics are obtained, such as Refusal rate, Drift rate, Attack Success rate (ASR) and Domain Vulnerability Index (DVI) which can allow a cross-domain comparison of safety behaviour in a systematic manner. It was empirically evaluated on 300 adversarial prompts per domain in which responses generated by a large language model were evaluated using automated scoring. The findings demonstrate significant differences in domain-specific safety behaviour. The medical domain demonstrates the lowest explicit harmful responses with the lowest Attack Success Rate, though it exhibits the highest drift rate through extensive contextual engagement and the lowest refusal rate. The law domain shows the lowest overall vulnerability as measured by the Domain Vulnerability Index, with moderate drift and explicit failures. In contrast, the finance domain emerges with the highest Attack Success Rate, indicating substantial susceptibility to adversarial prompts in professional financial advisory settings. Domains involving narrative framing or technical reasoning, particularly media and cybersecurity, exhibit the highest levels of safety boundary erosion. The cybersecurity domain demonstrates the highest Domain Vulnerability Index among all evaluated domains, while the media domain also shows significant vulnerability. These results suggest that the contextual framing of prompts plays a very important role in the way the models interpret and react to potentially unsafe requests. The data structure and assessment model suggested would offer a well-organised and scholarly suitable standard by which the behaviour of safety boundaries will be examined in domain-specific settings. The present work emphasizes the significance of domain-sensitive safety assessment and provides novel information regarding the weaknesses of contemporary language models by concentrating on indirect adversarial prompts in the context of realistic professional settings.

Article Details

Section
Articles