システム障害やRCAに関連した論文

アンケート

サーベイ

障害の解析 > OSINT

  • H. S. Gunawi, M. Hao, R. O. Suminto, A. Laksono, A. D. Satria, J. Adityatama, and K. J. Eliazar, “Why Does the Cloud Stop Computing?,” in Proceedings of the Seventh ACM Symposium on Cloud Computing, 2016, doi: 10.1145/2987550.2987583.
    • インターネットから収集した事例を解析して分析している論文
    • ダウンタイムのCDFや連鎖障害の組み合わせが紹介されている.
  • J. Sillito and E. Kutomi, “Failures and Fixes: A Study of Software System Incident Response,” in Proc. 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME), 2020, doi: 10.1109/icsme46990.2020.00027.
    • 公開されているポストモーテムをもとに障害を調査
  • X. Li, G. Yu, P. Chen, H. Chen, and Z. Chen, “Going through the Life Cycle of Faults in Clouds: Guidelines on Fault Handling,” in Proc. 2022 IEEE 33rd International Symposium on Software Reliability Engineering (ISSRE), 2022, doi: 10.1109/issre55969.2022.00022.
    • 公開されているポストモーテムをもとに障害を調査
    • Time to detectionやTime to recoveryのCDFがある.
    • 障害の原因調査のプロセスを整理した図がFig.2にある
    • コンフィグミスを解析している.
  • M. Ghanavati, D. Costa, J. Seboek, D. Lo, and A. Andrzejak, “Memory and resource leak defects and their repairs in Java projects,” Empirical Software Engineering, 2019, doi: 10.1007/s10664-019-09731-8.
    • JavaのOSSプロジェクト15個からメモリとリソースのリークを解析
  • M. Waseem, P. Liang, M. Shahin, A. Ahmad, and A. R. Nassab, “On the Nature of Issues in Five Open Source Microservices Systems: An Empirical Study,” in Proc. Evaluation and Assessment in Software Engineering, 2021, doi: 10.1145/3463274.3463337.
    • 5つのOSSのマイクロサービスシステムを対象に問題を解析した論文
      • 3つがMicroservice Framework
      • 2つがDemo Microservice Application
      • 評価対象のマイクロサービスがデモ用のものと,Microservice Frameworkであることには注意が必要.
      • 実際に運用されているアプリケーションのビジネスロジックに関する問題は含まれていないと思われる.
    • Figure 2の図で細かくツリーで問題をカテゴライズしているのが興味深い
    • 23.86%がTechnical Debt(技術的負債)で割合が最も多い
  • J. Sillito and M. Pope, “Failing and Learning: A Study of What is Learned About Reliability From Software Incidents,” in Proc. 2024 IEEE 35th International Symposium on Software Reliability Engineering Workshops (ISSREW), 2024, doi: 10.1109/issrew63542.2024.00093.
    • 89件の公開されているインシデントを調査したサーベイ
    • ISSREのワークショップで発表されている
    • インシデントの起因を集計しており,(1)コードのデプロイや有効化,(2)インフラ起因,(3)高負荷,(4)リソースリミットの超過,(5)メンテナンスが主要なトリガーになっていた.
  • Y. Zhao, L. Xiao, A. B. Bondi, B. Chen, and Y. Liu, “A Large-Scale Empirical Study of Real-Life Performance Issues in Open Source Projects,” IEEE Transactions on Software Engineering, 2023, doi: 10.1109/tse.2022.3167628.
    • GitHubにある13のOSSプロジェクトを解析してパフォーマンス問題を分析した論文
    • プログラミング言語(Python, C++, Java)によってパフォーマンス問題の根本原因の傾向に違いがある
    • Inefficient Data Structure (IDS)の分析が具体的な事例を含んでいる
  • Understanding and Dealing with Operator Mistakes in Internet Services | USENIX
    • 人間のオペレータを対象に作業の実験を行い,どんな作業ミスをするか分析している.

障害の解析 > 実システムの解析

  • Y. Chen, F. Ma, Y. Zhou, Z. Yan, Q. Liao, and Y. Jiang, “Themis: Finding Imbalance Failures in Distributed File Systems via a Load Variance Model,” in Proceedings of the Twentieth European Conference on Computer Systems, 2025, doi: 10.1145/3689031.3696082.
    • 清華大学のグループが分散ファイルシステムでのインバランス(不均衡)の障害を解析した論文
    • 実システムの障害を解析していて分散ファイルシステムならではのインバランスの障害に着目した論文はめずらしい
    • 53件の実システム(HDFSやCepfFS, GlusterFS, LeoFS)での障害を分析
  • Q. Xu, Y. Gao, and J. Wei, “An Empirical Study on Kubernetes Operator Bugs,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, 2024, doi: 10.1145/3650212.3680396.
    • Kubernetesのオペレータのバグを解析した論文.
    • とにかく色々なオペレータを解析している.
  • E. Kapel, L. Cruz, D. Spinellis, and A. Van Deursen, “On the Difficulty of Identifying Incident-Inducing Changes,” in Proceedings of the 46th International Conference on Software Engineering: Software Engineering in Practice, 2024, doi: 10.1145/3639477.3639755.
    • ソフトウェアの変更にともなう障害を分析している論文
    • 第一著者はオランダのING Bankの所属
  • Don’t Blame Disks for Every Storage Subsystem Failure
    • イリノイ大学とNetAppの共同研究でストレージシステムの障害を分析
  • S. Ghosh, M. Shetty, C. Bansal, and S. Nath, “How to fight production incidents?,” in Proceedings of the 13th Symposium on Cloud Computing, 2022, doi: 10.1145/3542929.3563482.
    • Microsoft Azureでの障害の分析をしている論文
    • 実際のインシデントの原因の内訳が書いてある.
    • インシデントの根本原因を分析すると27%がコードのバグ,依存による障害が16%,インフラストラクチャの障害が15%,データベースとネットワーク10%,Configのミスが12.5%だった.
  • L. A. Barroso, U. Hölzle, and P. Ranganathan, “The Datacenter as a Computer: Designing Warehouse-Scale Machines, Third Edition,” Synthesis Lectures on Computer Architecture, 2019, doi: 10.1007/978-3-031-01761-2.
    • Googleで障害原因の調査結果が紹介されている.
    • サービスレベルに影響を与える原因の1位はオペレータによるもの,2位はミスコンフィグレーションだった.
  • Z. Yin, X. Ma, J. Zheng, Y. Zhou, L. N. Bairavasundaram, and S. Pasupathy, “An empirical study on configuration errors in commercial and open source systems,” in Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles, 2011, doi: 10.1145/2043556.2043572.
    • NetAppの著者が含まれており,企業と大学での共同研究にみえる
    • ストレージベンダーへのテクニカルサポートへの問い合わせの内訳が書いてある.
    • 1位はハードウェア障害,2位は設定,3位は顧客環境,4位が顧客のナレッジ,5位がバグであった.
  • J. Chen, S. Zhang, X. He, Q. Lin, H. Zhang, D. Hao, Y. Kang, F. Gao, Z. Xu, Y. Dang, and D. Zhang, “How incidental are the incidents?,” in Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering, 2020, doi: 10.1145/3324884.3416624.
    • Microsoftで実際に発生したインシデントの分析をしている
    • 一時的なエラーが54%で最多だった.偽陽性アラートが16%で,設計上の問題が12%だった.
  • Y. Zhao, L. Jiang, Y. Tao, S. Zhang, C. Wu, Y. Wu, T. Jia, Y. Li, and Z. Wu, “How to Manage Change-Induced Incidents? Lessons from the Study of Incident Life Cycle,” in Proc. 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE), 2023, doi: 10.1109/issre59848.2023.00027.
    • Alibabaのインターンシップで学生がかいた論文
    • 231件の変更にともなう障害を定量的に解析している.
  • S. Cui, A. Patke, Z. Chen, A. Ranjan, H. Nguyen, P. Cao, B. Bode, G. Bauer, S. Jha, C. Narayanaswami, D. Sow, C. Di Martino, Z. T. Kalbarczyk, and R. K. Iyer, “Characterizing Modern GPU Resilience and Impact in HPC Systems: A Case Study of A100 GPUs,” in Proc. 2025 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W), 2025, doi: 10.1109/dsn-w65791.2025.00031.
    • NCSAのDelta(スーパーコンピュータ)の3年分の誤り回復データを分析した.
    • GPUメモリのGPUハードウェアよりもMTBE(Mean Time Between Errors)に関して160倍信頼性が高いことを示した.
  • Availability in Globally Distributed Storage Systems
    • Googleのグローバルに分散したストレージシステムの可用性を分析している.
  • Why Do Internet Services Fail, and What Can Be Done About It?
    • 3つのオンラインシステムの実データをもとに障害を解析している.3つの名前は不明.
    • Online: 成熟したオンラインサービス/インターネットポータル
    • Content: 最先端のグローバル・コンテンツ・ホスティング・サービス
    • ReadMostly: 成熟した読み取り(Read)主体のインターネットサービス
  • Y. Chen, H. Xie, M. Ma, Y. Kang, X. Gao, L. Shi, Y. Cao, X. Gao, H. Fan, M. Wen, J. Zeng, S. Ghosh, X. Zhang, C. Zhang, Q. Lin, S. Rajmohan, D. Zhang, and T. Xu, “Automatic Root Cause Analysis via Large Language Models for Cloud Incidents,” in Proceedings of the Nineteenth European Conference on Computer Systems, 2024, doi: 10.1145/3627703.3629553.
    • UIUC, Microsoft, et al.でLLMを使ってインシデントを解析
    • 単一データソースでは根本原因に届かない
    • 新規の根本原因が約25%
    • 再発するインシデントのうち大半(93.80%)は20日以内

障害の解析 > 新たな障害の発見・故障のモデリング

障害の解析 > HPCでの障害

障害の解析 > データセンタでの障害・ハードウェアの故障

障害の解析 > データセンタでの障害・ハードウェアの故障 > 故障の予測

障害の解析 > ネットワークでの障害

  • R. Potharaju and N. Jain, “When the network crumbles,” in Proceedings of the 4th annual Symposium on Cloud Computing, 2013, doi: 10.1145/2523616.2523638.
    • Microsoftのデータセンターネットワークの障害を解析している.
    • データセンター間とデータセンター内の2種類を解析している.
  • P. Gill, N. Jain, and N. Nagappan, “Understanding network failures in data centers,” in Proceedings of the ACM SIGCOMM 2011 conference, 2011, doi: 10.1145/2018436.2018477.
    • Microsoftのデータセンタのネットワークでの障害をLB,TRUNK,MGMT,CORE,ISC,IXごとに分析している.
    • Device problem typesとLink problem typesで障害の割合を分析している.
  • B. Yang, H. Hu, Y. Li, Y. Li, X. Tang, B. Tian, G. Wu, J. Xu, X. Zhang, F. Chen, C. Wang, E. Zhai, Y. Liao, D. Cai, and T. Lin, “SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud Infrastructures,” in Proceedings of the ACM SIGCOMM 2025 Conference, 2025, doi: 10.1145/3718958.3750536.
    • Alibabaでの具体的なデータセンターネットワーク障害の内訳を分類している
  • J. Meza, T. Xu, K. Veeraraghavan, and O. Mutlu, “A Large Scale Study of Data Center Network Reliability,” in Proceedings of the Internet Measurement Conference 2018, 2018, doi: 10.1145/3278532.3278566.
    • Facebookのデータセンターでのネットワークの障害を解析している.
    • データセンター間とデータセンター内の2種類を解析している.
  • P. Notaro, Q. Yu, S. Haeri, J. Cardoso, and M. Gerndt, “An Optical Transceiver Reliability Study based on SFP Monitoring and OS-level Metric Data,” in Proc. 2023 IEEE/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 2023, doi: 10.1109/ccgrid57682.2023.00011.
    • データセンター内のネットワークの光トランシーバの監視データとOSレベルメトリクスを使い,トランシーバ故障率と監視属性の正常動作範囲を推定
    • 著者の所属はHuaweiだった.

障害の解析 > 並列・並行システムのバグ

  • T. Leesatapornwongsa, J. F. Lukman, S. Lu, and H. S. Gunawi, “TaxDC,” in Proceedings of the Twenty-First International Conference on Architectural Support for Programming Languages and Operating Systems, 2016, doi: 10.1145/2872362.2872374.
    • データセンタで動作する分散システムの並列処理・並行処理のバグを分析している
  • Understanding, Detecting and Localizing Partial Failures in Large System Software | USENIX
    • ZookeeperやCassandraをはじめとした並列・並行システムのバグや故障を分析している.
  • A. Rabkin and R. H. Katz, “How Hadoop Clusters Break,” IEEE Software, 2013, doi: 10.1109/ms.2012.73.
    • Hadoopクラスタの障害の原因を分析している.
    • Flume, HBase, Zookeeperといったコンポーネントごとのエラー原因を分類している.
    • 下層の問題の原因(NFS, Java, Kerberos, DNS, Firewall)を分類している.
    • ミスコンフィグの内訳(RAM allocation, Thread allocation, Permissions, Other resources, Other, Absent/Malformed)を分類している.
    • 著者はClouderaで働いているプリンストン大学のポスドクだった

Misconfiguration / 設定ミス

  • H. Liu, S. Lu, M. Musuvathi, and S. Nath, “What bugs cause production cloud incidents?,” in Proceedings of the Workshop on Hot Topics in Operating Systems, 2019, doi: 10.1145/3317550.3321438.
    • Microsoft Azureの本番環境を対象に具体的なバグの種類を分類している.
    • 具体的にはバグを,データ形式のバグ,故障に関連したバグ,タイミングのバグ,定数のバグ,その他に分類している.
  • T. Xu, J. Zhang, P. Huang, J. Zheng, T. Sheng, D. Yuan, Y. Zhou, and S. Pasupathy, “Do not blame users for misconfigurations,” in Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles, 2013, doi: 10.1145/2517349.2522727.
    • Misconfigurationに着目したConfig errorを検知する論文
    • NetAppの著者が含まれており,企業と大学での共同研究にみえる
  • R. Bhagwan, S. Mehta, A. Radhakrishna, and S. Garg, “Learning Patterns in Configuration,” in Proc. 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE), 2021, doi: 10.1109/ase51524.2021.9678525.
    • 設定ミスの異常検知を行う論文
    • Microsoft Researchから出されている.

Configuration Management

  • C. Tang, T. Kooburat, P. Venkatachalam, A. Chander, Z. Wen, A. Narayanan, P. Dowell, and R. Karl, “Holistic configuration management at Facebook,” in Proceedings of the 25th Symposium on Operating Systems Principles, 2015, doi: 10.1145/2815400.2815401.
    • Facebookのコンフィグ管理システム

Production Microservice Analysis / マイクロサービスのプロダクション環境の分析

Anomaly detection / 異常検知

シングルモーダルRCA/シングルソースRCA > メトリクス

シングルモーダルRCA/シングルソースRCA > トレース

シングルモーダルRCA/シングルソースRCA > ログ

  • S. He, Q. Lin, J. G. Lou, H. Zhang, M. R. Lyu, and D. Zhang, “Identifying Impactful Service System Problems via Log Analysis,” in Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2018, doi: 10.1145/3236024.3236083.
  • Structured Comparative Analysis of Systems Logs to Diagnose Performance Problems
  • X. Zhang, Y. Xu, S. Qin, S. He, B. Qiao, Z. Li, H. Zhang, X. Li, Y. Dang, Q. Lin, M. Chintalapati, S. Rajmohan, and D. Zhang, “Onion: identifying incident-indicating logs for cloud systems,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2021, doi: 10.1145/3468264.3473919.
  • Q. Lin, H. Zhang, J. G. Lou, Y. Zhang, and X. Chen, “Log Clustering Based Problem Identification for Online Service Systems,” in Proceedings of the 38th International Conference on Software Engineering Companion, 2016, doi: 10.1145/2889160.2889232.
    • ログ件数の大規模化
      • “A Microsoft service system even generates over 1PB of logs every day.”
    • キーワード検索の限界(killやfailはダイナミックなインフラではfalse positiveになりやすい)
      • “The systems could proactively kill a job and restart it elsewhere, which causes many “kill” and “fail” keywords in logs.”
    • 再発した問題がすぐに解消されずに残ったままになるので,同じエラーログが前から出ていたままになっている.
      • “However, in a large-scale online service system, there are many recurrent issues, which could lead to a lot of redundant effort in examining logs and diagnosing the previously known problems.”
  • S. He, Q. Lin, J. G. Lou, H. Zhang, M. R. Lyu, and D. Zhang, “Identifying impactful service system problems via log analysis,” in Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2018, doi: 10.1145/3236024.3236083.
    • ログのクラスタリング手法を提案
    • Microsoftのプロダクション環境で実験を行った

マルチモーダルRCA/マルチソースRCA

Collection / 計装

グルーピング・要約 > Alert grouping / アラートグルーピング

グルーピング・要約 > Alert summary / アラートの要約

  • P. Jin, S. Zhang, M. Ma, H. Li, Y. Kang, L. Li, Y. Liu, B. Qiao, C. Zhang, P. Zhao, S. He, F. Sarro, Y. Dang, S. Rajmohan, Q. Lin, and D. Zhang, “Assess and Summarize: Improve Outage Understanding with Large Language Models,” in Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2023, doi: 10.1145/3611643.3613891.
    • LLMを使いインシデントのサマリーを行っている.
    • Microsoftで実際に3年にわたり運用してきたシステムであるそう.

グルーピング・要約 > Incident Linking / Ticket grouping

AIOps / Automation / Workflow

Auto-scaling

ログ > ログ検索エンジン / Log Search Engine

  • M. Yu, Z. Lin, J. Sun, R. Zhou, G. Jiang, H. Huang, and S. Zhang, “TencentCLS: the cloud log service with high query performances: Proceedings of the VLDB Endowment: Vol 15, No 12,” Proceedings of the VLDB Endowment, 2022, doi: 10.14778/3554821.3554837.
    • Tencentのログ管理のプラットフォームについて説明している.
    • 扱うログは,1日あたりペタバイトの規模が想定されている.
    • Apache Lucene 6.0でBKD Treeが導入されたがBKDツリーの複雑さは線形に相関があることを課題している.
  • W. Cao, X. Feng, B. Liang, T. Zhang, Y. Gao, Y. Zhang, and F. Li, “LogStore,” in Proceedings of the 2021 International Conference on Management of Data, 2021, doi: 10.1145/3448016.3457565.
    • Alibabaのログ管理プラットフォームを紹介している.
    • ヘビーな書き込みのスループットがあり,1秒あたり数千万のログレコードが書き込まれるという.
    • 検索では数十万に及ぶテナントがあり,ペタベイトに及ぶログを探すという.
    • Cost-effectiveなスケーラビリティのあるログストレージの設計が簡単でないことを課題としている.
  • B. Debnath, M. Solaimani, M. A. G. Gulzar, N. Arora, C. Lumezanu, J. Xu, B. Zong, H. Zhang, G. Jiang, and L. Khan, “LogLens: A Real-Time Log Analysis System,” in Proc. 2018 IEEE 38th International Conference on Distributed Computing Systems (ICDCS), 2018, doi: 10.1109/icdcs.2018.00105.
    • NEC Laboratories Americaの研究者が中心で執筆している.
    • リアルタイムのログ分析システムを提案した.
    • また,教師なし機械学習を使いアプリケーションログのパースを行った.
    • こうした,ログから異常なイベントを発見する方法や,ログメッセージのパーサーのパターンを自動で作成する方法は一つの研究テーマになっている印象がある.
  • T. Li, Y. Jiang, C. Zeng, B. Xia, Z. Liu, W. Zhou, X. Zhu, W. Wang, L. Zhang, J. Wu, L. Xue, and D. Bao, “FLAP,” in Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2017, doi: 10.1145/3097983.3098022.
    • フロリダ国際大学の研究者が中心で執筆している.
    • FIU Log Analysis Platformというイベントログを解析するためのプラットフォームで使われている技術を紹介している.
    • Challanges(課題)として以下の3つを主張している.
      • 多様な種類のイベントログが与えられるとき,どのようにイベント分析を広く一般的な方法でサポートするか.
      • 目的の異なる多様な分析の要件がある際に,どのように効率的に既存の分析手法を適用するか.
      • 多様な分析結果がある場合,どう効果的にユーザーへ提示するか.
  • H. Abe, K. Shima, D. Miyamoto, Y. Sekiya, T. Ishihara, K. Okada, R. Nakamura, and S. Matsuura, “Distributed Hayabusa,” in Proceedings of the Asian Internet Engineering Conference on - AINTEC ’19, 2019, doi: 10.1145/3340422.3343636.
    • 筆頭著者は日本のLepidum社(現在はGMO Cybersecurity by Ierae社)の方だった.
    • 共著者に東大の方が多い.
    • 大規模なログの検索のために複雑なストレージシステムやクラスタシステムを管理する必要があることを課題としていた.
    • Distributed Hayabusaというログ検索エンジンを提案している.
    • ログをタイムスタンプでSQLiteファイルに分割(シャーディング)することで高速化していた.
  • Read as Needed: Building WiSER, a Flash-Optimized Search Engine | USENIX
    • 検索エンジン WiSER を提案している.少ないメインメモリを使って高いスループットと低いレイテンシを出す手法を紹介している.
    • 以下を特徴として提案している.
      • データ配置の最適化
      • 2つのコストに配慮したブルームフィルター(特にここが新しそう)
      • 適応性のあるプリフェッチ
      • 容量と時間のトレードオフ
    • cf. ログ検索システムの論文まとめ | koyama’s blog

ログ > ログ件数の削減 / Log Volume Reduction

  • G. Yu, P. Chen, P. Li, T. Weng, H. Zheng, Y. Deng, and Z. Zheng, “LogReducer: Identify and Reduce Log Hotspots in Kernel on the Fly,” in Proc. 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), 2023, doi: 10.1109/ICSE48619.2023.00151.
    • eBPFを使ってオンライン・オフラインプロセスでログの件数を削減している.
    • 同一のテンプレートから繰り返し類似したログが出力されており,これがホットスポットになっている.
    • WeChatのプロダクションシステムで検証を行った.

評価 > Benchmark

評価 > Chaos Test