Exploiting Context-Aware Event Data for Fault Analysis
Abstract
Fault analysis in communication networks and distributed systems is a difficult process that heavily depends on system administrator’s experience and supporting tools. This process usually requires analytic techniques and several types of event data including log events, debug messages, trace obtained from these systems to investigate the root cause of faults. This paper introduces an approach of exploiting context-aware data and classification technique for improving this process. This approach uses both event data and context-aware data including CPU load, memory, processes, temperature, status to train a decision tree, and then applies the tree to assess suspected events. We have implemented and experimented the approach on the OpenStack cloud computing system with the Hadoop computing service and MELA event collection system. The experimental results reveal that the accuracy score of the approach reaches 85% on average. The paper also includes detailed analysis for the results.
Published
2016-10-10
Section
Regular articles
An author's submission implies that the manuscript has not been published previously, and is not currently submitted for publication elsewhere. Submission also implies that the Corresponding Author has consent of all authors (the Authors). Upon acceptance for publication transfer of copyright will be made to the Publisher of REV-JEC, who guarantees that full content of the published article is freely distributed on the Journal's website. The copyright transfer gives the Publisher of REV-JEC full authority to resolve any complaints of misuse or abuse (such as infringement or plagiarism) of the published article. The Authors have the freedom to redistribute and reuse the published article in any medium or format for any purpose, provided the original published article is properly cited. An article submission implies author agreement with this policy.