What Shadow AI Reveals About Workplace AI Adoption and Governance
A professional summary of Timothy Van Prooyen’s master’s thesis, with implications for IT, security, HR, and management
Generative AI has created an unusual governance problem. Employees can obtain sophisticated writing, coding, research, and analysis capabilities with little more than a browser and a personal account. Organizations, meanwhile, are still deciding which tools to approve, what data may be used, how outputs should be reviewed, and who is accountable when something goes wrong.
The result is Shadow AI: the use of AI systems, models, services, or APIs outside an organization’s approved governance structure. This can include an unapproved public tool, but it can also include an approved enterprise tool used for a task, data type, or decision that was never authorized.
Timothy Van Prooyen’s 2026 master’s thesis, Shadow AI in the Workplace: Balancing Productivity Gains and Organizational Risk, examines why employees use these tools despite known security, compliance, and accuracy risks. Its most important conclusion is not that employees are unaware of risk. It is that immediate, visible usefulness generally has more influence on adoption than risks that feel delayed, uncertain, or primarily organizational. That distinction matters because it changes the governance question from “How do we make employees understand the rules?” to “How do usefulness, workflow pressure, tool availability, confidence, and policy combine to shape actual behavior?” [1]
What the literature says about Shadow AI
Shadow AI grows out of a familiar workplace pattern
Shadow AI is new, but the underlying behavior is not. Research on Shadow IT has long shown that employees adopt spreadsheets, cloud services, messaging apps, and other unofficial tools when centrally provided systems do not meet changing work requirements. These workarounds can introduce duplication, weak controls, and security risk, but they can also be useful responses to slow provisioning, missing features, or rigid processes.
AI changes the scale of the issue. Older Shadow IT was often narrow and deterministic: an unofficial spreadsheet performed a known calculation or a cloud application handled a specific workflow. A general-purpose AI service can accept source code, contracts, employee information, strategy documents, or customer data; generate content that appears authoritative; and influence decisions in many different functions. Public tools may also sit outside enterprise identity, retention, logging, and monitoring systems. The governance problem is therefore not only an unauthorized application. It is the introduction of an externally hosted and partly opaque decision-support capability into everyday work. [1]
That difference has measurable security implications. IBM and the Ponemon Institute’s 2025 study of 600 breached organizations found that 20% reported a breach involving Shadow AI. Organizations with high levels of Shadow AI averaged $670,000 more in breach costs than organizations with little or no Shadow AI. These figures describe organizations in a breach study, not the probability that any organization will suffer a Shadow AI breach, but they illustrate why visibility, access control, and data governance matter. [2]
Three theories explain different parts of the behavior
The thesis combines three established theories rather than treating Shadow AI as a purely technical or compliance problem.
| Theoretical lens | Question it helps answer | Relevance to professionals |
|---|---|---|
| Technology Acceptance Model (TAM) | Does the employee believe the tool is useful and easy to use? | Explains why capable, low-friction tools attract experimentation and continued use. |
| Protection Motivation Theory (PMT) | How serious and likely does the risk feel, and does the employee believe it can be managed? | Explains why recognizing a threat does not automatically prevent use. |
| Principal–Agent Theory (PAT) | Do the employee and the organization carry the same incentives, information, and consequences? | Explains why a locally rational productivity choice can create organization-wide risk. |
TAM starts with two perceptions: usefulness and ease of use. These are beliefs, not independent measures of actual productivity or usability. Employees are likely to try a tool when it appears easy and continue when it helps them complete meaningful work. Natural-language interfaces have reduced the effort required to experiment so dramatically that an employee can test a new service before IT completes an initial vendor review.
The literature also describes a skill-leveling effect: AI assistance can produce particularly large gains for less experienced workers on some tasks. When a tool improves a draft, explains unfamiliar code, accelerates analysis, or helps someone perform closer to a more experienced colleague, continued use can feel professionally important rather than merely convenient.
PMT adds a complication. People evaluate both a threat and their ability to cope with it. An employee may know that confidential data should not enter a public model but believe that careful prompting, redaction, or human review makes the use safe. Whether that confidence is objectively justified is separate from its effect on behavior.
PAT places that decision inside the organization. Employees are rewarded for speed, output, and problem solving; the organization bears broader and often delayed consequences involving privacy, intellectual property, regulation, security, and reputation. No bad motive is required. The employee receives the immediate benefit while much of the downside is distributed elsewhere.
Together, the theories suggest that Shadow AI can be a rational response to the way work, incentives, and technology are configured. A reasonable policy can still lose influence when the approved path is unavailable, unclear, or poorly suited to the task.
Repeated workarounds can become normal work
The literature review uses normalization of deviance to explain how policy exceptions become routine. If an employee uses an outside tool, obtains a useful result, and observes no immediate harm, the behavior becomes easier to repeat. Coworkers share prompts or recommend tools, and an unofficial practice can become part of normal work while the formal policy remains unchanged.
Awareness and compliance are therefore different variables. A person can know a rule, understand some risk, and still decide that the practical benefit outweighs the expected cost. Governance is also likely to affect an occasional experimenter differently from someone whose daily process already depends on the tool.
Agentic AI raises the consequence of getting governance wrong
The literature distinguishes assistive generative AI from agentic AI. A chatbot usually produces an output for review. An agent may plan steps, retrieve information, call tools, modify files, browse sites, or initiate external actions. The concern therefore expands from content to authority, identity, permissions, and action.
OWASP highlights risks including goal manipulation, prompt injection, excessive permissions, unexpected tool use, and failures that propagate across connected systems. A system that can act should not be governed only as a system that can answer. [3]
How the study examined the problem
The thesis used an anonymous, cross-sectional survey of 317 knowledge workers recruited primarily through LinkedIn and SurveySwap. The survey was offered in English, German, and Spanish. Analyses requiring job classification used 273 respondents: 82 in IT-related roles and 191 in non-IT roles. The study tested five hypotheses covering usefulness, ease of use, tool disparity, policy awareness, role-based risk perception, self-efficacy, and the perceived risk of agentic AI.
The analysis included regression, structural equation modeling (SEM), group comparisons, and a supplementary machine-learning interpretation using SHAP. Forty optional comments added work context. The regressions and SEM examined specified relationships; SHAP offered an exploratory view of variable importance and possible non-linear patterns. [1]
What the results show
The best-fitting SEM included all 317 respondents and reported good fit. Its results are best read as a pattern rather than a ranking of isolated coefficients.
| Finding | Evidence from the thesis | Professional interpretation |
|---|---|---|
| Perceived usefulness was the strongest predictor of intention | Standardized β = .375, p < .001 | Employees are most attracted to unofficial AI when they believe it improves their work. |
| Ease of use predicted intention | β = .243, p < .001 | Low friction encourages trial and willingness to use a tool. |
| Ease of use did not independently predict usage in a separate regression | p = .495 after usefulness and other factors were included | Ease may open the door, but usefulness is more important to sustained use. |
| Perceived tool disparity had a small direct relationship with intention | β = .094, p = .044 | Approved-tool fit matters, but the effect was modest. |
| Tool disparity did not strengthen the usefulness relationship as hypothesized | PU × disparity, p = .807 in the SEM | Shadow AI cannot be explained only by saying the approved tool is worse. Habit, access, workflow fit, and other conditions may also matter. |
| Policy awareness was associated with lower usage overall | β = −.150, p = .009 | Clear policy has some value and should not be dismissed. |
| Policy awareness was not a significant deterrent among high-frequency users | p = .324 in the subgroup model | Communication alone may have little effect after unofficial use becomes embedded in a routine. |
| General perceived risk was not significant in the best SEM | β = −.063, p = .284 | Broad risk perception was a weaker and less stable influence than usefulness. |
| Innovativeness and self-efficacy predicted usage | β = .162 and .129, respectively | Confident experimenters are more likely to become active users, including outside approved channels. |
Usefulness was consistent; risk was more selective
Perceived usefulness was the most stable result. It strongly predicted intention, remained important in models of actual usage, and ranked first in the supplementary SHAP analysis.
The risk result needs more care. Risk was negatively associated with usage in one regression, but it did not significantly predict intention and was not significant in the best SEM. SHAP suggested that specific problems—especially confidential-data exposure and legal or compliance consequences—were more informative than describing unofficial AI as generally “risky.”
For security and training teams, “use AI responsibly” offers little help at the decision point. “Do not enter customer records into a public model,” “verify generated citations,” and “require approval before an agent sends an external message” connect risk to an observable task and response.
Policy matters, but its influence varies by user
The policy result is more balanced than either “policies work” or “policies are useless.” Greater awareness was associated with modestly lower use overall but was not a significant deterrent among high-frequency users.
For employees who are experimenting or uncertain, clear guidance may prevent unsafe adoption. For established users, the organization may need to understand the task, provide a credible alternative, change the workflow, or formalize a safer version of the practice.
The expected IT versus non-IT risk gap did not appear
IT respondents reported slightly higher average perceived risk than non-IT respondents—4.25 versus 3.96 on a seven-point scale—but the difference was not statistically significant (p = .115). IT respondents reported significantly greater ease of use, and their higher Shadow AI usage was marginal rather than conclusive at the broad role-group level. Software and AI developers, however, reported significantly higher usage than other roles (p = .017).
The proposed explanation that IT employees would perceive greater risk yet use Shadow AI more because of higher self-efficacy was not supported. “IT versus everyone else” may simply be too broad for governance: a developer testing code, a recruiter drafting a job description, and an analyst summarizing customer data face different capability needs and risks.
Respondents did not clearly distinguish agentic risk
The study found no significant difference between perceived risk for general generative AI and agentic AI, and no meaningful role-based difference. This does not establish equal objective risk: the survey used a limited measure and did not provide a detailed agent definition or task scenario.
The result may indicate a literacy gap. An assistant that drafts an email and an agent that can send it may look similar while creating very different control requirements. Systems should be classified by what they can access and do, not by how risky the interface feels.
What the findings mean for professionals
For IT and security, the quality of the approved path is part of the control environment. Blocking may be appropriate for high-risk uses, but it does not remove the task that led the employee to the tool. Intake, secure workspaces, identity controls, logging, and data protections work best when people can still accomplish the job.
For HR and learning teams, AI literacy should be role- and task-specific. The EU AI Act’s Article 4 calls for a sufficient level of literacy among relevant staff, with attention to knowledge, experience, context, and affected people—a stronger standard than one annual course for everyone. [4]
For managers, incentives deserve as much attention as awareness. Rewarding faster delivery while the approved process adds weeks of delay creates substitution pressure. It should also be safe to disclose an AI mistake, question an output, or request a better tool; psychological safety helps the organization learn before problems become incidents. [5]
For governance leaders, measurement should extend beyond licenses and blocked domains. Telemetry rarely explains the task, why a tool was chosen, whether the approved option was adequate, or whether the output created value. A complete view combines privacy-respecting feedback, usage, security events, workflow outcomes, quality, and incidents.
A practical governance agenda
The thesis recommends adaptive, behavior-aware governance rather than relying mainly on prohibition. A practical operating model can be organized around seven actions:
-
Map the work before selecting the control. Identify the task, data, decision, required capability, and consequence of error. Govern the use case, not only the product name.
-
Restrict, replace, or regularize the practice. Restrict unacceptable or high-consequence uses, replace risky workarounds with credible alternatives, and regularize useful lower-risk practices through documented tools and boundaries.
-
Reduce unnecessary tool disparity. Approved AI need not win every feature comparison, but it must cover needed capabilities with tolerable friction. An inaccessible license, an unusable model, or an unclear policy is not effective coverage.
-
Use task-based risk tiers. Enable low-risk work with guardrails, require approved tools and review for moderate-risk work, and restrict or require approval for confidential data, consequential decisions, regulated outputs, and external actions.
-
Make advanced users visible partners. Tool-intake channels, champion programs, and sandboxes can turn knowledgeable experimenters into a source of evidence rather than leaving their work outside governance.
-
Increase controls with autonomy and authority. Agents need least-privilege access, scoped identities, approval checkpoints, monitoring, audit trails, and reliable stop or recovery mechanisms. [3] [6]
-
Treat governance as a cycle. NIST’s Govern, Map, Measure, and Manage functions translate into maintaining an inventory, assessing higher-risk uses, monitoring behavior, recording incidents and exceptions, and revisiting controls as work changes. [7]
Responsibility remains distributed: IT and security manage identity, data, architecture, and controls; HR and learning teams build relevant skills; managers clarify acceptable use and incentives; legal, privacy, and risk teams review obligations and consequential uses; employees contribute essential workflow knowledge.
How to read the evidence
The study is exploratory. It used self-reported data from a digitally engaged convenience sample, so it should not be treated as a prevalence estimate for the entire workforce. The cross-sectional design identifies associations but cannot prove direction or causation; employees may use AI because they see it as useful, but frequent use may also increase perceived usefulness. Job-role data were missing for some respondents, the ease-of-use scale had low internal reliability, and the agentic-risk measure was limited.
Those limitations do not erase the pattern. They define its proper use. The thesis provides a strong basis for organizational diagnosis and field testing: identify which tasks create outside-tool pressure, determine where policy is unclear, compare actual and approved capabilities, test targeted interventions, and measure whether behavior and outcomes change. The most valuable next step is not to assume every organization behaves like the study sample. It is to measure how these forces interact inside a specific organization.
References
- Van Prooyen, T. (2026). Shadow AI in the Workplace: Balancing Productivity Gains and Organizational Risk. Prague University of Economics and Business.
- IBM & Ponemon Institute (2025). Cost of a Data Breach Report 2025: The AI Oversight Gap.
- OWASP Agentic Security Initiative (2025). Agentic AI—Threats and Mitigations.
- European Commission. AI Literacy—Questions & Answers.
- Edmondson, A. (1999). “Psychological Safety and Learning Behavior in Work Teams.” Administrative Science Quarterly, 44(2), 350–383.
- National Institute of Standards and Technology (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1).
- National Institute of Standards and Technology. AI Risk Management Framework Playbook.
From research to action
Behaviture AI Adoption Pulse turns these research questions into a privacy-first organizational diagnostic. Employees receive private, practical guidance, while leaders receive aggregated insight into adoption patterns, approved-tool fit, policy clarity, training needs, Shadow AI pressure, and readiness for agentic workflows—plus prioritized actions that can be measured in a future pulse. It helps IT, HR, security, and management move from assumptions about AI use to evidence-based decisions about enablement and governance.