TY - GEN
T1 - Temporal difference learning to detect unsafe system states
AU - Ning, Huazhong
AU - Xu, Wei
AU - Zhou, Yue
AU - Gong, Yihong
AU - Huang, Thomas
PY - 2008
Y1 - 2008
N2 - This paper proposes a general framework to detect unsafe states of a system whose basic realtime parameters are captured by multi-sensors. Our approach is to learn a danger level function which can be used to alert the users in advance of dangerous situations. The main challenge to this learning problem is the labelling issue, i.e., it is difficult to assign an objective danger level at each time step to the training data, except at the collapse points where a penalty can be assigned and at the successful ends where a certain reward can be assigned. In this paper, we treat the danger level as expected future reward (penalty is regarded as negative reward) and use temporal difference (TD) learning [2] to learn a function to approximate the expected future reward. The TD learning obtains the approximation by propagating the penalty/reward observable at collapse points or successful ends to the entire feature space following some constraints. Our approach is applied to, but not limited to, the application of monitoring of driving safety and the experimental results demonstrate the effectiveness of the approach.
AB - This paper proposes a general framework to detect unsafe states of a system whose basic realtime parameters are captured by multi-sensors. Our approach is to learn a danger level function which can be used to alert the users in advance of dangerous situations. The main challenge to this learning problem is the labelling issue, i.e., it is difficult to assign an objective danger level at each time step to the training data, except at the collapse points where a penalty can be assigned and at the successful ends where a certain reward can be assigned. In this paper, we treat the danger level as expected future reward (penalty is regarded as negative reward) and use temporal difference (TD) learning [2] to learn a function to approximate the expected future reward. The TD learning obtains the approximation by propagating the penalty/reward observable at collapse points or successful ends to the entire feature space following some constraints. Our approach is applied to, but not limited to, the application of monitoring of driving safety and the experimental results demonstrate the effectiveness of the approach.
KW - Driving safety
KW - Multi-sensor
KW - Temporal difference learning
KW - Unsafe system state
UR - https://www.scopus.com/pages/publications/77957969904
UR - https://www.scopus.com/pages/publications/77957969904#tab=citedBy
M3 - Conference contribution
AN - SCOPUS:77957969904
SN - 9781424421756
T3 - Proceedings - International Conference on Pattern Recognition
BT - 2008 19th International Conference on Pattern Recognition, ICPR 2008
T2 - 2008 19th International Conference on Pattern Recognition, ICPR 2008
Y2 - 8 December 2008 through 11 December 2008
ER -