H. S. Gunawi, M. Hao, R. O. Suminto, A. Laksono, A. D. Satria, J. Adityatama, and K. J. Eliazar, “Why Does the Cloud Stop Computing?,” in Proceedings of the Seventh ACM Symposium on Cloud Computing, 2016, doi: 10.1145/2987550.2987583.
Q. Xu, Y. Gao, and J. Wei, “An Empirical Study on Kubernetes Operator Bugs,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, 2024, doi: 10.1145/3650212.3680396.
Kubernetesのオペレータのバグを解析した論文.
とにかく色々なオペレータを解析している.
E. Kapel, L. Cruz, D. Spinellis, and A. Van Deursen, “On the Difficulty of Identifying Incident-Inducing Changes,” in Proceedings of the 46th International Conference on Software Engineering: Software Engineering in Practice, 2024, doi: 10.1145/3639477.3639755.
S. Ghosh, M. Shetty, C. Bansal, and S. Nath, “How to fight production incidents?,” in Proceedings of the 13th Symposium on Cloud Computing, 2022, doi: 10.1145/3542929.3563482.
J. Chen, S. Zhang, X. He, Q. Lin, H. Zhang, D. Hao, Y. Kang, F. Gao, Z. Xu, Y. Dang, and D. Zhang, “How incidental are the incidents?,” in Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering, 2020, doi: 10.1145/3324884.3416624.
S. Cui, A. Patke, Z. Chen, A. Ranjan, H. Nguyen, P. Cao, B. Bode, G. Bauer, S. Jha, C. Narayanaswami, D. Sow, C. Di Martino, Z. T. Kalbarczyk, and R. K. Iyer, “Characterizing Modern GPU Resilience and Impact in HPC Systems: A Case Study of A100 GPUs,” in Proc. 2025 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W), 2025, doi: 10.1109/dsn-w65791.2025.00031.
NCSAのDelta(スーパーコンピュータ)の3年分の誤り回復データを分析した.
GPUメモリのGPUハードウェアよりもMTBE(Mean Time Between Errors)に関して160倍信頼性が高いことを示した.
Y. Chen, H. Xie, M. Ma, Y. Kang, X. Gao, L. Shi, Y. Cao, X. Gao, H. Fan, M. Wen, J. Zeng, S. Ghosh, X. Zhang, C. Zhang, Q. Lin, S. Rajmohan, D. Zhang, and T. Xu, “Automatic Root Cause Analysis via Large Language Models for Cloud Incidents,” in Proceedings of the Nineteenth European Conference on Computer Systems, 2024, doi: 10.1145/3627703.3629553.
UIUC, Microsoft, et al.でLLMを使ってインシデントを解析
単一データソースでは根本原因に届かない
新規の根本原因が約25%
再発するインシデントのうち大半(93.80%)は20日以内
障害の解析 > 新たな障害の発見・故障のモデリング
P. Huang, C. Guo, L. Zhou, J. R. Lorch, Y. Dang, M. Chintalapati, and R. Yao, “Gray Failure,” in Proceedings of the 16th Workshop on Hot Topics in Operating Systems, 2017, doi: 10.1145/3102980.3103005.
Microsoft Azureで実際に発生したGray Failure(部分的な故障)について紹介している論文
N. Bronson, A. Aghayev, A. Charapko, and T. Zhu, “Metastable failures in distributed systems,” in Proceedings of the Workshop on Hot Topics in Operating Systems, 2021, doi: 10.1145/3458336.3465286.
A. Oliner and J. Stearley, “What Supercomputers Say: A Study of Five System Logs,” in Proc. 37th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN′07), 2007, doi: 10.1109/dsn.2007.103.
S. Cui, A. Patke, Z. Chen, A. Ranjan, H. Nguyen, P. Cao, B. Bode, G. Bauer, S. Jha, C. Narayanaswami, D. Sow, C. Di Martino, Z. T. Kalbarczyk, and R. K. Iyer, “Characterizing Modern GPU Resilience and Impact in HPC Systems: A Case Study of A100 GPUs,” in Proc. 2025 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W), 2025, doi: 10.1109/dsn-w65791.2025.00031.
NCSAのDelta(スーパーコンピュータ)の3年分の誤り回復データを分析した.
GPUメモリのGPUハードウェアよりもMTBE(Mean Time Between Errors)に関して160倍信頼性が高いことを示した.
F. Yu, H. Xu, S. Jian, C. Huang, Y. Wang, and Z. Wu, “DRAM Failure Prediction in Large-Scale Data Centers,” in Proc. 2021 IEEE International Conference on Joint Cloud Computing (JCC), 2021, doi: 10.1109/jcc53141.2021.00012.
J. Alter, J. Xue, A. Dimnaku, and E. Smirni, “SSD failures in the field,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2019, doi: 10.1145/3295500.3356172.
GoogleのデータセンターでのSSDの故障を分析,予測している.
障害の解析 > ネットワークでの障害
R. Potharaju and N. Jain, “When the network crumbles,” in Proceedings of the 4th annual Symposium on Cloud Computing, 2013, doi: 10.1145/2523616.2523638.
T. Leesatapornwongsa, J. F. Lukman, S. Lu, and H. S. Gunawi, “TaxDC,” in Proceedings of the Twenty-First International Conference on Architectural Support for Programming Languages and Operating Systems, 2016, doi: 10.1145/2872362.2872374.
ミスコンフィグの内訳(RAM allocation, Thread allocation, Permissions, Other resources, Other, Absent/Malformed)を分類している.
著者はClouderaで働いているプリンストン大学のポスドクだった
Misconfiguration / 設定ミス
H. Liu, S. Lu, M. Musuvathi, and S. Nath, “What bugs cause production cloud incidents?,” in Proceedings of the Workshop on Hot Topics in Operating Systems, 2019, doi: 10.1145/3317550.3321438.
T. Xu, J. Zhang, P. Huang, J. Zheng, T. Sheng, D. Yuan, Y. Zhou, and S. Pasupathy, “Do not blame users for misconfigurations,” in Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles, 2013, doi: 10.1145/2517349.2522727.
Misconfigurationに着目したConfig errorを検知する論文
NetAppの著者が含まれており,企業と大学での共同研究にみえる
R. Bhagwan, S. Mehta, A. Radhakrishna, and S. Garg, “Learning Patterns in Configuration,” in Proc. 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE), 2021, doi: 10.1109/ase51524.2021.9678525.
設定ミスの異常検知を行う論文
Microsoft Researchから出されている.
Configuration Management
C. Tang, T. Kooburat, P. Venkatachalam, A. Chander, Z. Wen, A. Narayanan, P. Dowell, and R. Karl, “Holistic configuration management at Facebook,” in Proceedings of the 25th Symposium on Operating Systems Principles, 2015, doi: 10.1145/2815400.2815401.
Facebookのコンフィグ管理システム
Production Microservice Analysis / マイクロサービスのプロダクション環境の分析
Y. Gan, Y. Zhang, D. Cheng, A. Shetty, P. Rathi, N. Katarki, A. Bruno, J. Hu, B. Ritchken, B. Jackson, K. Hu, M. Pancholi, Y. He, B. Clancy, C. Colen, F. Wen, C. Leung, S. Wang, L. Zaruvinsky, M. Espinosa, R. Lin, Z. Liu, J. Padilla, and C. Delimitrou, “Unveiling the Hardware and Software Implications of Microservices in Cloud and Edge Systems,” IEEE Micro, 2020, doi: 10.1109/mm.2020.2985960.
具体的な商用マイクロサービスの規模感(Netflix)を説明している記事
S. Luo, H. Xu, C. Lu, K. Ye, G. Xu, L. Zhang, Y. Ding, J. He, and C. Xu, “Characterizing Microservice Dependency and Performance,” in Proceedings of the ACM Symposium on Cloud Computing, 2021, doi: 10.1145/3472883.3487003.
I. T. A. Lee, Z. Zhang, A. Parwal, and M. Chabbi, “The Tale of Errors in Microservices,” Proceedings of the ACM on Measurement and Analysis of Computing Systems, 2024, doi: 10.1145/3700436.
M. Barletta, M. Cinque, C. Di Martino, Z. T. Kalbarczyk, and R. K. Iyer, “Mutiny! How Does Kubernetes Fail, and What Can We Do About It?,” in Proc. 2024 54th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN), 2024, doi: 10.1109/dsn58291.2024.00016.
Kubernetesに関して実際に起きた障害をまとめている.
Misconfigurationがシステム全体の障害に伝搬していることを説明している.
“Both experiments and real-world K8s failures show that one incorrect data value can propagate and cause system-wide failures despite the resiliency strategies.”
K. Seemakhupt, B. E. Stephens, S. Khan, S. Liu, H. Wassel, S. H. Yeganeh, A. C. Snoeren, A. Krishnamurthy, D. E. Culler, and H. M. Levy, “A Cloud-Scale Characterization of Remote Procedure Calls,” in Proceedings of the 29th Symposium on Operating Systems Principles, 2023, doi: 10.1145/3600006.3613156.
P. Srinivas, F. Husain, A. Parayil, A. Choure, C. Bansal, and S. Rajmohan, “Intelligent Monitoring Framework for Cloud Services: A Data-Driven Approach,” in Proceedings of the 46th International Conference on Software Engineering: Software Engineering in Practice, 2024, doi: 10.1145/3639477.3639753.
Z. Li, J. Chen, R. Jiao, N. Zhao, Z. Wang, S. Zhang, Y. Wu, L. Jiang, L. Yan, Z. Wang, Z. Chen, W. Zhang, X. Nie, K. Sui, and D. Pei, “Practical Root Cause Localization for Microservice Systems via Trace Analysis,” in Proc. 2021 IEEE/ACM 29th International Symposium on Quality of Service (IWQOS), 2021, doi: 10.1109/iwqos52092.2021.9521340.
MicrosoftのエンジニアがマイクロサービスベースのシステムでのRoot Cause Analysisのために,冗長なグラフを構造を取り除く手法を提案
シングルモーダルRCA/シングルソースRCA > ログ
S. He, Q. Lin, J. G. Lou, H. Zhang, M. R. Lyu, and D. Zhang, “Identifying Impactful Service System Problems via Log Analysis,” in Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2018, doi: 10.1145/3236024.3236083.
X. Zhang, Y. Xu, S. Qin, S. He, B. Qiao, Z. Li, H. Zhang, X. Li, Y. Dang, Q. Lin, M. Chintalapati, S. Rajmohan, and D. Zhang, “Onion: identifying incident-indicating logs for cloud systems,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2021, doi: 10.1145/3468264.3473919.
“However, in a large-scale online service system, there are many recurrent issues, which could lead to a lot of redundant effort in examining logs and diagnosing the previously known problems.”
S. He, Q. Lin, J. G. Lou, H. Zhang, M. R. Lyu, and D. Zhang, “Identifying impactful service system problems via log analysis,” in Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2018, doi: 10.1145/3236024.3236083.
Y. Wang, G. Li, Z. Wang, Y. Kang, Y. Zhou, H. Zhang, F. Gao, J. Sun, L. Yang, P. Lee, Z. Xu, P. Zhao, B. Qiao, L. Li, X. Zhang, and Q. Lin, “Fast Outage Analysis of Large-Scale Production Clouds with Service Correlation Mining,” in Proc. 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE), 2021, doi: 10.1109/icse43902.2021.00085.
J. Mace, R. Roelke, and R. Fonseca, “Pivot tracing,” in Proceedings of the 25th Symposium on Operating Systems Principles, 2015, doi: 10.1145/2815400.2815415.
N. Zhao, J. Chen, X. Peng, H. Wang, X. Wu, Y. Zhang, Z. Chen, X. Zheng, X. Nie, G. Wang, Y. Wu, F. Zhou, W. Zhang, K. Sui, and D. Pei, “Understanding and Handling Alert Storm for Online Service Systems,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering: Software Engineering in Practice, 2020, doi: 10.1145/3377813.3381363.
Microsoftで障害の検知を自動化を行っている.障害の検知までにかかる時間(Time To Detect)が紹介されている.
V. Ganatra, A. Parayil, S. Ghosh, Y. Kang, M. Ma, C. Bansal, S. Nath, and J. Mace, “Detection Is Better Than Cure: A Cloud Incidents Perspective,” in Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2023, doi: 10.1145/3611643.3613898.
P. Jin, S. Zhang, M. Ma, H. Li, Y. Kang, L. Li, Y. Liu, B. Qiao, C. Zhang, P. Zhao, S. He, F. Sarro, Y. Dang, S. Rajmohan, Q. Lin, and D. Zhang, “Assess and Summarize: Improve Outage Understanding with Large Language Models,” in Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2023, doi: 10.1145/3611643.3613891.
LLMを使いインシデントのサマリーを行っている.
Microsoftで実際に3年にわたり運用してきたシステムであるそう.
グルーピング・要約 > Incident Linking / Ticket grouping
S. Ghosh, K. Grover, J. Wong, C. Bansal, R. Namineni, M. Verma, and S. Rajmohan, “Dependency Aware Incident Linking in Large Cloud Systems,” in Proc. Companion Proceedings of the ACM Web Conference 2024, 2024, doi: 10.1145/3589335.3648311.
J. Liu, S. He, Z. Chen, L. Li, Y. Kang, X. Zhang, P. He, H. Zhang, Q. Lin, Z. Xu, S. Rajmohan, D. Zhang, and M. R. Lyu, “Incident-aware Duplicate Ticket Aggregation for Cloud Systems,” in Proc. 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), 2023, doi: 10.1109/icse48619.2023.00193.
AIOps / Automation / Workflow
M. Shetty, C. Bansal, S. P. Upadhyayula, A. Radhakrishna, and A. Gupta, “AutoTSG: learning and synthesis for incident troubleshooting,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2022, doi: 10.1145/3540250.3558958.
K. Rzadca, P. Findeisen, J. Swiderski, P. Zych, P. Broniek, J. Kusmierek, P. Nowak, B. Strack, P. Witusowski, S. Hand, and J. Wilkes, “Autopilot,” in Proceedings of the Fifteenth European Conference on Computer Systems, 2020, doi: 10.1145/3342195.3387524.
W. Cao, X. Feng, B. Liang, T. Zhang, Y. Gao, Y. Zhang, and F. Li, “LogStore,” in Proceedings of the 2021 International Conference on Management of Data, 2021, doi: 10.1145/3448016.3457565.
B. Debnath, M. Solaimani, M. A. G. Gulzar, N. Arora, C. Lumezanu, J. Xu, B. Zong, H. Zhang, G. Jiang, and L. Khan, “LogLens: A Real-Time Log Analysis System,” in Proc. 2018 IEEE 38th International Conference on Distributed Computing Systems (ICDCS), 2018, doi: 10.1109/icdcs.2018.00105.
T. Li, Y. Jiang, C. Zeng, B. Xia, Z. Liu, W. Zhou, X. Zhu, W. Wang, L. Zhang, J. Wu, L. Xue, and D. Bao, “FLAP,” in Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2017, doi: 10.1145/3097983.3098022.
H. Abe, K. Shima, D. Miyamoto, Y. Sekiya, T. Ishihara, K. Okada, R. Nakamura, and S. Matsuura, “Distributed Hayabusa,” in Proceedings of the Asian Internet Engineering Conference on - AINTEC ’19, 2019, doi: 10.1145/3340422.3343636.
筆頭著者は日本のLepidum社(現在はGMO Cybersecurity by Ierae社)の方だった.
RMIT University + The University of Newcastle + University of New South Wales
S. Jha, R. Arora, Y. Watanabe, T. Yanagawa, Y. Chen, J. Clark, B. Bhavya, M. Verma, H. Kumar, H. Kitahara, N. Zheutlin, S. Takano, D. Pathak, F. George, X. Wu, B. O. Turkkan, G. Vanloo, M. Nidd, T. Dai, O. Chatterjee, P. Gupta, S. Samanta, P. Aggarwal, R. Lee, P. Murali, J. W. Ahn, D. Kar, A. Rahane, C. Fonseca, A. Paradkar, Y. Deng, P. Moogi, P. Mohapatra, N. Abe, C. Narayanaswami, T. Xu, L. R. Varshney, R. Mahindru, A. Sailer, L. Shwartz, D. Sow, N. C. M. Fuller, and R. Puri, “ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks,” arXiv:2502.05352, 2025.
V. Heorhiadi, S. Rajagopalan, H. Jamjoom, M. K. Reiter, and V. Sekar, “Gremlin: Systematic Resilience Testing of Microservices,” in Proc. 2016 IEEE 36th International Conference on Distributed Computing Systems (ICDCS), 2016, doi: 10.1109/icdcs.2016.11.
A. Basiri, L. Hochstein, N. Jones, and H. Tucker, “Automating chaos experiments in production,” in Proc. 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 2019, doi: 10.1109/ICSE-SEIP.2019.00012.