Skip to main content

07. References & Standards

This directory lists authoritative benchmarks, regulatory frameworks, foundational academic literature, and open-source toolkits for Machine Learning Security and Adversarial AI engineering.


1. Security Standards & Risk Taxonomies​


2. Landmark Academic Research Papers​

Adversarial Evasion & Robustness​

  • Goodfellow, I. J., Shlens, J., & Szegedy, C. (2014). Explaining and Harnessing Adversarial Examples. arXiv preprint arXiv:1412.6572.
    (Introduced the Fast Gradient Sign Method (FGSM) and linear model hypothesis).
  • Madry, A., Makelov, A., Schmidt, L., Tsipras, D., & Vladu, A. (2017). Towards Deep Learning Models Resistant to Adversarial Attacks. arXiv preprint arXiv:1706.06083.
    (Introduced PGD iterative attack and minimax adversarial training formulation).
  • Carlini, N., & Wagner, D. (2017). Towards Evaluating the Robustness of Neural Networks. IEEE Symposium on Security and Privacy (SP).
    (Formulated C&W L2/L_inf optimization attacks bypassing defensive distillation).

Data Poisoning & Trojan Backdoors​

  • Gu, T., Dolan-Gavitt, B., & Garg, S. (2017). BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain. IEEE Access.
    (First paper demonstrating backdoor Trojan trigger injection in neural network weights).
  • Shafahi, A., Huang, W. R., Najibi, M., Suciu, O., Studer, C., Dumitras, T., & Goldstein, T. (2018). Poison Frogs! Clean-Label Backdoor Attacks on Machine Learning. NeurIPS.
    (Demonstrated targeted clean-label feature collision data poisoning without altering training labels).
  • Tran, B., Li, J., & Madry, A. (2018). Spectral Signatures in Backdoor Attacks. NeurIPS.
    (Introduced SVD feature representation filtering to detect and remove backdoor training samples).

Model Intellectual Property Theft & Privacy Leakage​

  • TramΓ¨r, F., Zhang, F., Juels, A., Reiter, M. K., & Ristenpart, T. (2016). Stealing Machine Learning Models via Prediction APIs. USENIX Security Symposium.
    (Pioneered black-box model extraction attacks against commercial prediction APIs).
  • Shokri, R., Stronati, M., Song, C., & Shmatikov, V. (2017). Membership Inference Attacks Against Machine Learning Models. IEEE Symposium on Security and Privacy (SP).
    (Formulated shadow model techniques to infer individual record membership in training sets).

3. Open-Source Security Toolkits & Frameworks​

Tool / RepositoryOrganizationLinkPurpose
Adversarial Robustness Toolbox (ART)Linux FoundationGitHubComprehensive python library for ML security evaluation and defense.
Microsoft CounterfitMicrosoft AzureGitHubAutomation tool for black-box penetration testing of ML models.
SafetensorsHugging FaceGitHubSafe, zero-copy tensor serialization format replacing unsafe pickle files.
PicklescanHugging FaceGitHubSecurity scanner for detecting dangerous opcodes in Python pickle artifacts.
ModelscanProtect AIGitHubOpen-source tool for scanning ML model artifacts for code execution risks.
FicklingTrail of BitsGitHubDecompiler and static analyzer for inspecting Python pickle bytecode.
PyTorch OpacusMeta AIGitHubHigh-speed library for training PyTorch models with Differential Privacy (DP-SGD).
RobustBenchCommunityRobustBench LeaderboardStandardized benchmark leaderboard for empirical adversarial robustness evaluation.

Return to Module Overview & Index β†’