AutoClassWeb (v2.2.1) is a web interface to AutoClass C, an unsupervised Bayesian classification system developped by the NASA.

It utilizes AutoClassWrapper (v1.5.1), a Python wrapper for AutoClass C.


Input data

Input data file is tab-delimited and must comply with the Tab-separated values format.

The first line must be a header with column names. Avoid accentuated or special characters (e.g.: $, &, !, /, β) or space. These characters will be automatically replaced by _. Also avoid lengthy column names. Column names must be unique.

The first column must be gene/protein/object names.

Missing data are allowed and must be encoded with an empty value, i.e. nothing, (don't use NA, ?, None or ' ').
Example (missing value for prot1 in exp2 and prot3 in exp1):

name	exp1	exp2	exp3
prot1	0.123		550.61
prot2	0.003	4.966	27.77
prot3		7.723	9345.34

Your data must belong to one of the three following categories:

Real Location

Negative and positive real values (e.g.: microarray log ratio, elevation, position).
Example:

name	exp4	exp5	exp6
prot1	-18.3	1.723	-5.6151
prot2	14.7	-0.006	-2.7779
prot3	-22.5	0.023	9.3441

Real Scalar

Singly bounded real values, typically bounded below at zero (e.g.: length, weight, age).
Example:

name	exp1	exp2	exp3
prot1	0.123	1.723	550.61
prot2	0.003	4.966	27.77
prot3	1.812	7.723	9345.34

Discrete

Qualitative data (e.g.: color, phenotype, name...).
Example:

name	tissu	condition
prot1	muscle	light
prot2	nervous	light
prot3	muscle	shadow

Note: if your dataset contains several types of data (real scalar, real location, discrete), split your dataset into multiple datasets with homogeneous data type.

Classification job

Classification jobs are listed in the status page.

The maximum running time for a job is 480 hours. Results older than 30 days are automatically deleted.

Upon successful classification, results are available through a download link in the status page.

Results

Results are bundled in a zip archive with the following files:

  • something_out.cdt and something_out_withproba.cdt can be open with Java TreeView. The file something_out_withproba.cdt exhibits the probability for each gene/protein/object to belong to each class.
  • something_out_stats.tsv contain means and standard deviations of numeric columns (real scalar and real location) for each class.
  • something_out_dendrogram.png is a dendrogram plot representing the distance between all classes.
  • something_out.tsv contains all the data with the class assignement and membership probabilities for all classes. This file is in the Tab-separated values format and can then easily be parsed with Excel, R, Python...

More on the AutoClass algorithm

AutoClass is an unsupervised Bayesian classification system. It offers multiple advantages:

  • Number of classes are determined automatically.
  • Missing values are allowed.
  • Real and discrete values can be combined (in separate input files).
  • Class membership probabilities are also computed.

For more information, see the NASA documentation on AutoClass C.