A Max-Relevance-Min-Divergence criterion for data discretization with applications on naive Bayes

Shihe Wang; Jianfeng Ren; Ruibin Bai; Yuan Yao; Xudong Jiang

doi:10.1016/j.patcog.2023.110236

A Max-Relevance-Min-Divergence criterion for data discretization with applications on naive Bayes

Shihe Wang, Jianfeng Ren, Ruibin Bai, Yuan Yao, Xudong Jiang

School of Computer Science

Research output: Journal Publication › Article › peer-review

6 Citations (Scopus)

Abstract

In many classification models, data is discretized to better estimate its distribution. Existing discretization methods often target at maximizing the discriminant power of discretized data, while overlooking the fact that the primary target of data discretization in classification is to improve the generalization performance. As a result, the data tend to be over-split into many small bins since the data without discretization retain the maximal discriminant information. Thus, we propose a Max-Dependency-Min-Divergence (MDmD) criterion that maximizes both the discriminant information and generalization ability of the discretized data. More specifically, the Max-Dependency criterion maximizes the statistical dependency between the discretized data and the classification variable while the Min-Divergence criterion explicitly minimizes the JS-divergence between the training data and the validation data for a given discretization scheme. The proposed MDmD criterion is technically appealing, but it is difficult to reliably estimate the high-order joint distributions of attributes and the classification variable. We hence further propose a more practical solution, Max-Relevance-Min-Divergence (MRmD) discretization scheme, where each attribute is discretized separately, by simultaneously maximizing the discriminant information and the generalization ability of the discretized data. The proposed MRmD is compared with the state-of-the-art discretization algorithms under the naive Bayes classification framework on 45 benchmark datasets. It significantly outperforms all the compared methods on most of the datasets.

Original language	English
Article number	110236
Journal	Pattern Recognition
Volume	149
DOIs	https://doi.org/10.1016/j.patcog.2023.110236
Publication status	Published - May 2024

Keywords

Data discretization
Maximal dependency
Maximal relevance
Minimal divergence
Naive Bayes classification

ASJC Scopus subject areas

Software
Signal Processing
Computer Vision and Pattern Recognition
Artificial Intelligence

Access to Document

10.1016/j.patcog.2023.110236

Cite this

@article{f3eb7e86681240b79b0646e2a9d3e679,

title = "A Max-Relevance-Min-Divergence criterion for data discretization with applications on naive Bayes",

abstract = "In many classification models, data is discretized to better estimate its distribution. Existing discretization methods often target at maximizing the discriminant power of discretized data, while overlooking the fact that the primary target of data discretization in classification is to improve the generalization performance. As a result, the data tend to be over-split into many small bins since the data without discretization retain the maximal discriminant information. Thus, we propose a Max-Dependency-Min-Divergence (MDmD) criterion that maximizes both the discriminant information and generalization ability of the discretized data. More specifically, the Max-Dependency criterion maximizes the statistical dependency between the discretized data and the classification variable while the Min-Divergence criterion explicitly minimizes the JS-divergence between the training data and the validation data for a given discretization scheme. The proposed MDmD criterion is technically appealing, but it is difficult to reliably estimate the high-order joint distributions of attributes and the classification variable. We hence further propose a more practical solution, Max-Relevance-Min-Divergence (MRmD) discretization scheme, where each attribute is discretized separately, by simultaneously maximizing the discriminant information and the generalization ability of the discretized data. The proposed MRmD is compared with the state-of-the-art discretization algorithms under the naive Bayes classification framework on 45 benchmark datasets. It significantly outperforms all the compared methods on most of the datasets.",

keywords = "Data discretization, Maximal dependency, Maximal relevance, Minimal divergence, Naive Bayes classification",

author = "Shihe Wang and Jianfeng Ren and Ruibin Bai and Yuan Yao and Xudong Jiang",

note = "Publisher Copyright: {\textcopyright} 2023 The Authors",

year = "2024",

month = may,

doi = "10.1016/j.patcog.2023.110236",

language = "English",

volume = "149",

journal = "Pattern Recognition",

issn = "0031-3203",

publisher = "Elsevier Ltd.",

}

TY - JOUR

T1 - A Max-Relevance-Min-Divergence criterion for data discretization with applications on naive Bayes

AU - Wang, Shihe

AU - Ren, Jianfeng

AU - Bai, Ruibin

AU - Yao, Yuan

AU - Jiang, Xudong

PY - 2024/5

Y1 - 2024/5

N2 - In many classification models, data is discretized to better estimate its distribution. Existing discretization methods often target at maximizing the discriminant power of discretized data, while overlooking the fact that the primary target of data discretization in classification is to improve the generalization performance. As a result, the data tend to be over-split into many small bins since the data without discretization retain the maximal discriminant information. Thus, we propose a Max-Dependency-Min-Divergence (MDmD) criterion that maximizes both the discriminant information and generalization ability of the discretized data. More specifically, the Max-Dependency criterion maximizes the statistical dependency between the discretized data and the classification variable while the Min-Divergence criterion explicitly minimizes the JS-divergence between the training data and the validation data for a given discretization scheme. The proposed MDmD criterion is technically appealing, but it is difficult to reliably estimate the high-order joint distributions of attributes and the classification variable. We hence further propose a more practical solution, Max-Relevance-Min-Divergence (MRmD) discretization scheme, where each attribute is discretized separately, by simultaneously maximizing the discriminant information and the generalization ability of the discretized data. The proposed MRmD is compared with the state-of-the-art discretization algorithms under the naive Bayes classification framework on 45 benchmark datasets. It significantly outperforms all the compared methods on most of the datasets.

AB - In many classification models, data is discretized to better estimate its distribution. Existing discretization methods often target at maximizing the discriminant power of discretized data, while overlooking the fact that the primary target of data discretization in classification is to improve the generalization performance. As a result, the data tend to be over-split into many small bins since the data without discretization retain the maximal discriminant information. Thus, we propose a Max-Dependency-Min-Divergence (MDmD) criterion that maximizes both the discriminant information and generalization ability of the discretized data. More specifically, the Max-Dependency criterion maximizes the statistical dependency between the discretized data and the classification variable while the Min-Divergence criterion explicitly minimizes the JS-divergence between the training data and the validation data for a given discretization scheme. The proposed MDmD criterion is technically appealing, but it is difficult to reliably estimate the high-order joint distributions of attributes and the classification variable. We hence further propose a more practical solution, Max-Relevance-Min-Divergence (MRmD) discretization scheme, where each attribute is discretized separately, by simultaneously maximizing the discriminant information and the generalization ability of the discretized data. The proposed MRmD is compared with the state-of-the-art discretization algorithms under the naive Bayes classification framework on 45 benchmark datasets. It significantly outperforms all the compared methods on most of the datasets.

KW - Data discretization

KW - Maximal dependency

KW - Maximal relevance

KW - Minimal divergence

KW - Naive Bayes classification

UR - http://www.scopus.com/inward/record.url?scp=85183464630&partnerID=8YFLogxK

U2 - 10.1016/j.patcog.2023.110236

DO - 10.1016/j.patcog.2023.110236

M3 - Article

AN - SCOPUS:85183464630

SN - 0031-3203

VL - 149

JO - Pattern Recognition

JF - Pattern Recognition

M1 - 110236

ER -

A Max-Relevance-Min-Divergence criterion for data discretization with applications on naive Bayes

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this