AutoClassWeb (v2.2.1) is a web interface to AutoClass C, an unsupervised Bayesian classification system developped by the NASA.
It utilizes AutoClassWrapper (v1.5.1), a Python wrapper for AutoClass C.
Input data
Input data file is tab-delimited and must comply with the Tab-separated values format.
The first line must be a header with column names.
Avoid accentuated or special characters (e.g.: $, &, !, /, β) or space.
These characters will be automatically replaced by _. Also avoid lengthy column names.
Column names must be unique.
The first column must be gene/protein/object names.
Missing data are allowed and must be encoded with an empty value, i.e. nothing,
(don't use NA, ?, None or ' ').
Example (missing value for prot1 in exp2
and prot3 in exp1):
name exp1 exp2 exp3 prot1 0.123 550.61 prot2 0.003 4.966 27.77 prot3 7.723 9345.34
Your data must belong to one of the three following categories:
Real Location
Negative and positive real values (e.g.: microarray log ratio, elevation, position).
Example:
name exp4 exp5 exp6 prot1 -18.3 1.723 -5.6151 prot2 14.7 -0.006 -2.7779 prot3 -22.5 0.023 9.3441
Real Scalar
Singly bounded real values, typically bounded below at zero
(e.g.: length, weight, age).
Example:
name exp1 exp2 exp3 prot1 0.123 1.723 550.61 prot2 0.003 4.966 27.77 prot3 1.812 7.723 9345.34
Discrete
Qualitative data (e.g.: color, phenotype, name...).
Example:
name tissu condition prot1 muscle light prot2 nervous light prot3 muscle shadow
Note: if your dataset contains several types of data (real scalar, real location, discrete), split your dataset into multiple datasets with homogeneous data type.
Classification job
Classification jobs are listed in the status page.
The maximum running time for a job is 480 hours. Results older than 30 days are automatically deleted.
Upon successful classification, results are available through a download link in the status page.
Results
Results are bundled in a zip archive with the following files:
-
something_out.cdtandsomething_out_withproba.cdtcan be open with Java TreeView. The filesomething_out_withproba.cdtexhibits the probability for each gene/protein/object to belong to each class. -
something_out_stats.tsvcontain means and standard deviations of numeric columns (real scalar andreal location ) for each class. -
something_out_dendrogram.pngis a dendrogram plot representing the distance between all classes. -
something_out.tsvcontains all the data with the class assignement and membership probabilities for all classes. This file is in the Tab-separated values format and can then easily be parsed with Excel, R, Python...
More on the AutoClass algorithm
AutoClass is an unsupervised Bayesian classification system. It offers multiple advantages:
- Number of classes are determined automatically.
- Missing values are allowed.
- Real and discrete values can be combined (in separate input files).
- Class membership probabilities are also computed.