This repository contains the code and data used to evaluate computational tools for enzyme function prediction. We benchmarked multiple EC (Enzyme Commission) number prediction algorithms — spanning both similarity-based and machine/deep learning approaches — using reaction SMILES as input. The pipeline includes scripts for dataset preprocessing, tool evaluation under various conditions, performance assessment across EC hierarchy levels and classes, and visualization of results. Overall, this repository provides a reproducible and extensible framework for benchmarking EC number prediction methods and helps users identify the most suitable tool for their metabolic modeling applications.
General pipeline of EC number prediction methodologies. A) Similarity-based methods. B) Machine/deep learning methods.
Specifically, we assessed the tools under three conditions:
- We evaluated all selected methods — E-zyme, E-zyme2, BridgIT, SelenzymeRF, SIMMER, Theia, BEC-Pred and CLAIRE — using 20% of the KEGG 2025 database (1866 reactions). We also evaluated these using a subset of 500 Rhea 2025 reactions.
- For all of the methods with available source code — SelenzymeRF, SIMMER, Theia, BEC-Pred and CLAIRE — we evaluated them using three different data splits of the Rhea 2025 dataset: Stratified random split, Time-based split and Scaffold-aware split. We trained the models with the training dataset and tested with the test set for each split. Additionally, we trained or used as prior knowledge 90% of the MetaNetX/KEGG/ECREACT/Rhea databases, and then queried the methods with the remaining 10%. All mentioned splits are included in
data. - We did a case study on 28 drugs and their associated enzyme-annotated degradation reactions, and used them to query against all selected methods. Additionally, we applied a Top1 and Top5 majority voting strategy using SelenzymeRF, SIMMER, Theia and BEC-Pred, to show the potential of combining multiple algorithms to correctly predict EC number.
For more information, please refer to:
- Josefina Arcagni: jarcagniriv@unav.es
- Telmo Blasco: tblasco@tecnun.es
The code has the following structure:
ECNumberPrediction/
├── data/ # Input datasets for analysis
├── methods/ # Implemented EC number prediction methods
│ ├── BEC-Pred
│ ├── BridgIT
│ ├── CLAIRE
│ ├── E-zyme
│ ├── SelenzymeRF
│ ├── SIMMER
│ └── theia
├── results/ # Output results from the analyses
│ ├── Case1
│ ├── Case2
│ ├── CaseStudy
│ └── MajorityVote
└── README.md
- Put input files under
data/. - Each method in
methods/has its own bash file with implementation steps. - Outputs for runs and evaluations go to
results/<method_or_case>/. - Example workflow:
- Prepare inputs in
data/. - Run a method from
methods/run_<MethodName>.sh/. - Check results in
results/.
- Prepare inputs in
Clone the repository:
git clone https://github.com/PlanesLab/ECNumberPrediction.git
cd ECNumberPredictionEach method in the methods/ folder may have its own installation requirements. Refer to the individual method documentation for setup instructions.
| Tool | Year | Type | Database | Features | Open-source code |
|---|---|---|---|---|---|
| E-zyme | 2009 | SB | KEGG | RDM patterns, substrate-product pairs, Tanimoto score | No |
| E-zyme2 | 2016 | SB | KEGG | RDM patterns, substrate-product pairs, graph-based substructures | No |
| BridgIT | 2019 | SB | KEGG | Daylight fingerprints, reactive site identification, BNICE.ch rules | No |
| SelenzymeRF | 2023 | SB | MetaNetX | Morgan fingerprints, RXNMapper reactive sites, fragment analysis | Yes (GitHub) |
| SIMMER | 2023 | SB | MetaCyc | Atom-Pair fingerprints, Tanimoto score, enrichment analysis | Yes (GitHub) |
| Theia | 2023 | ML | ECREACT / Rhea | MLP, differential reaction fingerprints | Yes (GitHub) |
| BEC-Pred | 2024 | ML | USPTO-ECREACT | BERT, transfer learning | Yes (GitHub) |
| CLAIRE | 2025 | ML | ECREACT | Contrastive learning, rxnfp embeddings, differential reaction fingerprints | Yes (GitHub) |
The table above summarizes the tools used, detailing their year of release, type (SB: similarity-based or ML: machine learning), associated databases, key features, and availability of open-source code.
Canonicalizes reaction SMILES using RDKit. Accepts a plain text file (one SMILES per line), a CSV/TSV file with a specified SMILES column, or an entire directory of such files.
Find in data/Preprocessing/canonicalize_rxn_SMILES.py
Each reaction SMILES is expected in the format:
reactant1.reactant2>>product1.product2
Atom mapping is not assumed. Reactants and products are each sorted alphabetically after canonicalization to ensure a deterministic output regardless of input order. Invalid molecules are filtered out; invalid reactions are skipped with a warning.
Plain text file (one SMILES per line):
python3 canonicalize_smiles.py \
--input data/reaction_smiles.txt \
--output data/reaction_smiles_can.txtCSV file with a named SMILES column:
python3 canonicalize_smiles.py \
--input data/reactions.csv \
--output data/reactions_can.csv \
--smiles_col rxn_smilesTSV file, custom output column name:
python3 canonicalize_smiles.py \
--input data/reactions.tsv \
--output data/reactions_can.tsv \
--smiles_col smiles \
--output_col can_smiles| Argument | Default | Description |
|---|---|---|
--input |
— | Input file or directory (required) |
--output |
— | Output file or directory (required) |
--smiles_col |
— | Column name containing SMILES (required for CSV/TSV) |
--output_col |
canonical_smiles |
Column name for the canonical SMILES output |
--sep |
auto | Delimiter for tabular files (auto-detected from .csv/.tsv extension) |
--keep_failed |
off | Keep rows where canonicalization failed (written as empty) |
--extension |
.txt |
File extension filter when processing a directory |
For plain text input, each valid reaction is written as a canonical SMILES string on its own line.
For CSV/TSV input, all original columns are preserved and a new column
(default: canonical_smiles) is appended with the canonicalized result.
Rows that fail canonicalization are dropped unless --keep_failed is set.
A summary of canonicalized vs. failed reactions is printed per file.
- Create & activate venv:
python3 -m venv e-zyme_venv
source e-zyme_venv/bin/activate - Install libs:
python -m pip install --upgrade pip
pip install requests beautifulsoup4 pandas- Make runner executable and run:
chmod +x methods/E-zyme/run_ezyme.sh
bash methods/E-zyme/run_ezyme.shThis will run the script for Case 1 and Case Study and the result CSV files for E-zyme and E-zyme2 will both be saved in the results/ folder under the corresponding case folders with the tool name: E-zyme.csv.
- Create & activate venv:
python3 -m venv bridgit_venv
source bridgit_venv/bin/activate - Install libs:
python -m pip install --upgrade pip
pip install requests pandas rdkit- Make runner executable and run:
chmod +x methods/BridgIT/run_BridgIT.sh
bash methods/BridgIT/run_BridgIT.shNotes: BridgIT can only be accessed through its own server in https://lcsb-databases.epfl.ch/Bridgit, a user account needs to be created to access it. The first part of the bash file will process reaction SMILES and create the necessary files in methods/BridgIT/input/ to input in their web server. Once the results are ready, download them from the server and place them in the methods/BridgIT/output/ folder. The steps for extracting the results are in the second part of the bash file.
This will run the script for Case 1 and Case Study and the result csv files will be saved in the results/ folder under the corresponding case folders with the tool name: BridgIT.csv.
- Create & activate venv:
python3 -m venv selenzyme_venv
source selenzyme_venv/bin/activate - Install libs:
python -m pip install --upgrade pip
pip install requests beautifulsoup4 pandas- Clone repository & download data:
git clone -b SelenzymeRF https://github.com/synbiochem/selenzyme.git methods/SelenzymeRF/SelenzymeRF_code
cd methods/SelenzymeRF/SelenzymeRF_codeNotes: unzip the required reference datasets as indicated in the upstream repository. The script already included in Selenzyme_code/: /methods/SelenzymeRF/SelenzymeRF_code/start_server.sh was specifically modified to run the cases and shouldn't be deleted.
- Make runner executable and run:
chmod +x /methods/SelenzymeRF/run_selenzymerf.sh
bash /methods/SelenzymeRF/run_selenzymerf.shThis will run the script for all cases and the result csv files will be saved in the results/ folder under the corresponding case folders with the tool name: SelenzymeRF.csv.
- Create conda environment:
conda env create -f methods/SIMMER/simmer_env.yml
conda activate simmer_env- Clone repository & download data:
git clone https://github.com/aebustion/SIMMER.git methods/SIMMER/SIMMER_code
cd methods/SIMMER/SIMMER_codeNotes: import the required reference datasets as indicated in the upstream repository. The scripts already included in SIMMER_code/: methods/SIMMER/SIMMER_code/SIMMER/SIMMER.py and methods/SIMMER/SIMMER_code/SIMMER/SIMMER2.py were specifically modified to run the cases and shouldn't be deleted.
- Make runner executable and run:
chmod +x methods/SIMMER/run_SIMMER.sh
bash methods/SIMMER/run_SIMMER.shThis will run the script for all cases and the result csv files will be saved in the results/ folder under the corresponding case folders with the tool name: SIMMER.csv.
Requirements: 1 GPU and 50G memory for training and evaluating the model.
- Create conda environment:
conda env create -f methods/theia/theia_env.yml
conda activate theia_env- Clone repository & download data:
git clone https://github.com/daenuprobst/theia.git methods/theia/theia_code
cd methods/theia/theia_codeNotes: All the scripts already included in Theia_code/ were specifically modified to run the cases and shouldn't be deleted.
- Make runner executable and run:
chmod +x methods/theia/run_theia.sh
bash methods/theia/run_theia.shThis will run the script for all cases and the result csv files will be saved in the results/ folder under the corresponding case folders with the tool name: theia.csv.
Requirements: 1 GPU and 128G memory for training and evaluating the model.
- Create conda environment:
conda env create -f methods/BEC-Pred/becpred_gpu.yml
conda activate becpred_env- Make runner executable and run:
chmod +x methods/BEC-Pred/run_becpred.sh
bash methods/BEC-Pred/run_becpred.shThis will run the script for all cases and the result csv files will be saved in the results/ folder under the corresponding case folders with the tool name: BEC-Pred.csv.
Requirements: 1 GPU and 128G memory for training and evaluating the model.
- Create conda environments:
conda env create -f methods/CLAIRE/claire_env.yml
# activate when using CLAIRE tools
conda activate claire_env
conda env create -f methods/CLAIRE/rxnfp_env.yml
# activate when using rxnfp utilities
conda activate rxnfp_env- Clone repository & download data:
git clone https://github.com/zishuozeng/CLAIRE.git methods/CLAIRE/CLAIRE_code
cd methods/CLAIRE/CLAIRE_codeNotes: Import or unzip any reference datasets required by CLAIRE as indicated in the upstream repository. Do not remove or modify the modified scripts included in methods/CLAIRE/CLAIRE_code/ that are needed to run the cases.
- Make runner executable and run:
chmod +x methods/CLAIRE/run_claire.sh
bash methods/CLAIRE/run_claire.shThis will run CLAIRE for the configured cases and save result CSV files under the corresponding case folders in results/ with the tool name: CLAIRE.csv.
Results are organized in subfolders inside results/:
-
results/Case1/– Results for the first evaluation case (queried all methods with their original dataset tested with KEGG reaction queries). -
results/Case2/– Results for the second evaluation case (five open-source methods trained on 80% of MetaNetX dataset and queried on the other 20%). -
results/CaseStudy/– Results for the Case Study (queries all methods with their original dataset using 28 drug-associated reactions). -
results/MajorityVote/– Top1 and Top5 majority voting strategies using SelenzymeRF, SIMMER, Theia and BEC-Pred.
To reproduce the metrics and figures used in the paper for each case, follow the steps below:
- Merge all tool CSV outputs into merged_output.csv:
python3 results/Case1/join_results.pyOutput: results/Case1/merged_output.csv
- Compute evaluation metrics:
python3 results/Case1/get_metrics.pyOutput: results/Case1/evaluation_summary.csv
- Generate plots:
Rscript results/Case1/metrics-case1.RThe plot will be saved in results/Case1/case1_plot.png.
- Merge all tool CSV outputs into merged_output.csv:
python3 results/Case2/join_results.pyOutput: results/Case2/merged_output.csv
- Compute evaluation metrics:
python3 results/Case2/get_metrics.pyOutput: results/Case2/evaluation_summary.csv
- Generate plots:
Rscript results/Case2/results-metanetx/metrics-case2.RThe plot will be saved in results/Case2/results-metanetx/case2_plot.jpg.
- Merge all tool CSV outputs into merged_output.csv:
python3 results/CaseStudy/join_results.pyOutput: results/CaseStudy/merged_output.csv
- Generate plots:
Rscript results/CaseStudy/case_study_heatmap.RThe plot will be saved in results/CaseStudy/casestudyplot.png.
A tool to combine multiple EC number prediction outputs using weighted majority voting across multiple prediction methods.
Each method's predictions are parsed into ranked groups. All ECs in the same group receive equal points based on their rank position:
- Rank 1 →
top_npts - Rank 2 →
top_n - 1pts - ...
- Rank N → 1 pt
Points are summed across all methods to produce a total weighted score per EC. Co-predictions (pipe-separated ECs at the same rank) receive equal points and are not penalised.
Tie-breaking (all descending): total weighted score → count at rank 1 → count at rank 2 → ... → if still tied, pipe-joined into one shared slot (e.g. 1.2.3|4.5.6).
ECs are automatically collapsed to the 3rd hierarchical level (e.g. 1.2.3.4 → 1.2.3).
A CSV file with an identifier column (default: entity) and one or more method columns containing EC predictions in this format:
1.1.1.1|1.1.1.2;2.7.1.1|2.7.1.2
;separates rank slots (rank 1 → rank 2 → ...)|separates tied ECs within the same rank slot
All entities:
python3 majority_vote.py \
--input_csv results/merged_predictions.csv \
--entity all \
--use_all \
--output_csv results/majority_votes.csvSingle entity:
python3 majority_vote.py \
--input_csv results/merged_predictions.csv \
--entity R001 \
--methods MethodA MethodB MethodCCustom identifier column or top-N window:
python3 majority_vote.py \
--input_csv results/merged_predictions.csv \
--entity all \
--use_all \
--id_col reaction_id \
--top_n 3| Argument | Default | Description |
|---|---|---|
--input_csv |
— | Path to merged prediction CSV (required) |
--entity |
— | Entity ID to query, or all (required) |
--methods |
— | Specific method columns to use |
--use_all |
— | Use all EC prediction columns |
--top_n |
5 |
Scoring window and output length |
--id_col |
entity |
Name of the identifier column |
--output_csv |
— | Optional path to save results CSV |
Either --methods or --use_all must be provided.
| Column | Description |
|---|---|
majority_top1 |
Top-ranked consensus EC (or A|B if unbreakable tie) |
majority_topN_ranked |
Full ranked list of up to N slots, separated by ; |