Understanding AI Threats and Attack Vectors
Overview of AI Model Security
- What makes AI models vulnerable
- AI security vs traditional software
security
- AI attack taxonomy: training time
vs inference time
Data Poisoning and Training Attacks
- Poisoned inputs and backdoors
- Label flipping and data injection
Adversarial Examples and Inference-Time Attacks
- How small changes can fool AI
- Image, text, and audio adversarial
attacks
Discussion and Case Studies
- Real-world examples (e.g.,
autonomous driving, facial recognition)
- Failures due to adversarial or
poisoned data
Model
Theft, Privacy, and Defense Mechanisms
Model Extraction and Inversion Attacks
- API-based model stealing
- Membership inference and model
inversion
- Risks of exposing model outputs and
confidence scores
Defense Strategies – Part 1
- Adversarial training and gradient
masking
- Model ensembling and input
preprocessing
- Hands-on: Apply basic defenses to a
model and test robustness
Privacy-Preserving Machine Learning
- Differential privacy
- Federated learning and secure
aggregation
- Use of tools like PySyft and Opacus
for privacy
Secure Deployment and Advanced Defenses
Secure Model Deployment and Access Control
- Threats in model hosting and
serving environments
- API throttling, output
sanitization, and rate limiting
- Securing endpoints and pipelines
AI Red Teaming and Monitoring Tools
- Overview of red teaming AI models
- Toolkits: IBM Adversarial
Robustness Toolbox (ART), CleverHans, Foolbox
- Monitoring AI systems for anomalies
Wrap-Up and Review
- Review of tools, techniques, and
best practices
- Discussion on evolving threats and
trends
- Q&A, resources, and course
feedback