Welcome to PON-P4

VUS-aware predictor for amino acid substitutions


PON-P4 predicts the pathogenicity of single amino acid substitutions in human MANE-specific proteins using a three-class prediction framework. The predictor classifies variants into Pathogenic, Benign, or Variant of Uncertain Significance (VUS).

The method is based on machine learning and integrates diverse biological information, including gene-level, protein-level, variation-specific, and structural features. PON-P4 was trained on a large curated dataset of experimentally and clinically characterized amino acid substitutions.

PON-P4 utilizes a gradient boosting framework optimized for high predictive performance, scalability, and rapid inference. The predictor has been extensively benchmarked against state-of-the-art pathogenicity prediction methods.


Development

PON-P4 is developed in the Protein Structure and Bioinformatics Research (LU-PSB) group led by Prof. Mauno Vihinen at Lund University, Sweden.

How to cite

We are preparing a manuscript describing the method. In the meantime, please cite the method with the url http://structure.bmc.lu.se/PON-P4/

Last updated: 2026-06-24

Predict protein variations

Format: refseq_ids,position,refAA,altAA

Note: Maximum 1000 lines.
One protein variation per line.
Example: NP_000008.1,104,K,N





Prediction results

Unpredicted protein substitutions

Protein prediction table columns:
  • Input: Original input coordinates in the format you provided
  • RefSeq protein: RefSeq protein identifier
  • Protein variation: Variation provided.
  • Predicted label: P = Pathogenic, B = Benign, V = VUS
  • Comments: Additional information or notes regarding the prediction
Note: Substitution are not predicted in the first and last protein position. Alteration in the start codon prevent protein synthesis, substitutions at the c-terminal resiude lead to insertions.

Predict transcript variations

Format: nm_id:c.positionREF>ALT

Note: Maximum 1000 lines.
One transcript variation per line.
Example: NM_000018.4:c.1726T>C





Prediction results

Unpredicted transcript substitutions

Transcript prediction table columns:
  • Input transcript variation: Original transcript-level variation
  • RefSeq protein: RefSeq protein identifier
  • Protein variation: Converted protein variation
  • Predicted label: P = Pathogenic, B = Benign, V = VUS
  • Comments: Additional information or notes
Note: Substitution are not predicted in the first and last protein position. Alteration in the start codon prevent protein synthesis, substitutions at the c-terminal resiude lead to insertions.

Predict genomic variations

Format: chrNUM:g.posREF>ALT

Note: Maximum 1000 lines.
One genomic variation per line.
Example: chr10:g.123456A>G





Prediction results

Unpredicted genomic substitutions

Genomic prediction table columns:
  • Input genomic variation: Original genomic-level variation
  • RefSeq protein: RefSeq protein identifier
  • Protein variation: Converted protein variation
  • Predicted label: P = Pathogenic, B = Benign, V = VUS
  • Comments: Additional information or notes
Note: Substitution are not predicted in the first and last protein position. Alteration in the start codon prevent protein synthesis, substitutions at the c-terminal resiude lead to insertions.

Select identifier

1. Choose identifier type
2. Enter ID

Obtain all possible predictions for a protein by providing its ID.

Download Heat Map (PDF)

Download Results (CSV)
Note: Precomputed results are based on MANE-selected (MANE v1.3) RefSeq protein identifiers.

Summary

Predicted substitution pathogenicity

Heatmap for predicted pathogenicity

1. About PON-P4

PON-P4 is a three-class predictor for amino acid substitution pathogenicity.

The platform offers three interfaces for processing data:

  • Protein variation
  • Transcript variation
  • Genomic variation

Each of these interfaces requires inputs in a specific format, and variations have to be described using MANE-compliant references.

Results are provided in both CSV and PDF formats.

Input protein variation

Use the text box where protein variations can be pasted or typed, one protein variation per line.

NP_000005.3,1292,V,F NP_000005.3,436,W,E NP_000008.1,382,Y,W NP_000008.1,412,S,N

Input transcript variation

Use the text box where transcript variations can be pasted or typed, one transcript variation per line.

NM_000016.6:c.613G>C NM_000016.6:c.653C>G NM_000017.4:c.319C>T

Input genomic variation

Use the text box where genomic variations can be pasted or typed, one genomic variation per line.

chr1:g.75745819G>C chr2:g.47789620T>A chr3:g.12156789A>G

2. Datasets

Datasets were obtained with extensive data mining from ClinVar and LOVD.

The training and test datasets are also available at VariBench Dataset 35.


3. Download files

  • Predictions for the tested methods. Download
  • Proteome-wide PON-P4 predictions for all possible substitutions in a single file: Download
  • PON-P4 predictions for 13 mitochondrial proteins: Download
  • List of HCG IDs excluded from MANE: Download
  • IDs or excluded proteins for which all necessary data are not available. Download

4. Citing PON-P4

https://structure.bmc.lu.se/PON-P4/


5. Contact

If you have any problems, please contact Muhammad Kabir at muhammad.kabir@med.lu.se or Prof. Mauno Vihinen at mauno.vihinen@med.lu.se.

1. Disclaimer

This non-profit server, its associated data and services are for research purposes only. The responsibility of Protein Structure and Bioinformatics Group, Lund University is limited to applying the best efforts in providing and publishing good programs and data. The developers have no responsibility for the usage of results, data or information which this server has provided.

2. Liability

In preparation of this site and service, every effort has been made to offer the most current and correct information possible. We render no warranty, express or implied, as to its accuracy or that the information is fit for a particular purpose, and will not be held responsible for any direct, indirect, putative, special, or consequential damages arising out of any inaccuracies or omissions. In utilizing this service, individuals, organizations, and companies absolve Lund University or any of their employees or agents from responsibility for the effect of any process, method or product that may be produced or adopted by any party, notwithstanding that the formulation of such process, method or product may be based upon information provided here.

3. License

The PON-P4 content is licensed under the Creative Commons Attribution-NonCommercial (CC BY-NC) license, allowing others to share and adapt the material for non-commercial purposes with appropriate credit.

4. Compliance with EU AI Act Systemic Risk under Article 55 (read complete article)

  1. 4.1 Conduct model evaluations and adversarial testing with the aim of identifying and mitigating systemic risks. PON-P4 has been extensively tested with 10-fold cross-validation and blind test sets, showing good performance in all tests.
  2. 4.2 Assess and mitigate possible systemic risks at the EU level caused by the development, deployment, or use of general-purpose AI models. PON-P4 can be used as one component in the clinical diagnosis of genetic diseases. According to ACMG guidelines, a single type of evidence cannot alone be used for diagnosis.
  3. 4.3 Track and report serious incidents to the AI Office and National Competent Authorities. Any potential incidents will be reported accordingly.
  4. 4.4 Ensure adequate cybersecurity protections for the model and its physical infrastructure. The server running the PON-P4 predictor is protected by Lund University Division of IT firewalls, and backups are performed regularly.

5. Compliance with the MINimum Information for Medical AI Reporting (MINIMAR) by (Tina et al. 2020) requirements

5.1 Study population and setting

  • Population: Not relevant.
  • Study setting: Method development based on clinically defined variations.
  • Data source: ClinVar and LOVD.
  • Cohort selection: Healthy and diseased individuals.

5.2 Patient demographic characteristics

Not relevant.

5.3 Model architecture

  • Model output: Prediction whether a variation is B, P or VUS.
  • Target user: Research or clinical laboratory.
  • Data splitting: Disjointed training and test datasets in 10:1 ratio.
  • Gold standard: Blind test set of variants classified in ClinVar or LOVD. All variants in a protein were either in the training or test set.
  • Model task: Classification.
  • Model architecture: First LightGBM gradient boosting classifier to predict pathogenic and benign variants. Further classification of the remaining cases with another LightGBM predictor.
  • Features: 31 selected features for the first predictor, 40 for the second. See Supplementary Table 14.
  • Missingness: None.

5.4 Model evaluation

  • Optimization: Selection of the algorithm, feature selection, and parameter optimisation with Optuna for the final predictor.
  • Internal model validation: 10-fold cross-validation.
  • External model validation: Blind test set.
  • Transparency: The data used for training and testing is freely available. The predictor is freely available. Predictions for all possible 19 substitutions are available for all MANE transcript-based proteins.

6. Compliance with the REFORMS checklist ( Kapoor et al. Sci Adv 2024 )

6.1 Study goals

  1. 1a. Population about which the scientific claim is made: We collected all classified variants from ClinVar and LOVD which fulfilled the inclusion criteria.
  2. 1b. Motivation for choosing this population: ClinVar and LOVD were the major sources of variation information.
  3. 1c. Motivation for the use of ML methods in the study: Variation interpretation is a complex problem in which numerous factors contribute to the variation effect, making it difficult or impossible to handle without ML approaches.

6.2 Computational reproducibility

  1. 2a. Dataset used for training and evaluating the model: A search using the keyword "missense" on the ClinVar website was used to identify amino acid substitutions. The results were filtered using "single nucleotide" as the variation type, "missense" as the molecular result, and "germline" as the classification type. The review status for both benign and pathogenic variations was set to two stars or above. The data were further filtered with proprietary scripts. Duplicates and cases with unknown variations, insertions, or deletions were excluded. Additional pathogenic variants were obtained from the LOVD shared databases using the defined inclusion and exclusion criteria. Duplicates between datasets were eliminated. The datasets are available at the VariBench database, dataset 35: VariBench dataset 35
  2. 2b. Code used to train and evaluate the model: Python 3.8.18 was used together with standard scientific toolchains. LightGBM served as the primary gradient boosting engine. scikit-learn was used to wrap the base LightGBM classifier in an ECOC multiclass strategy via OutputCodeClassifier and to partition datasets with KFold. Optuna, incorporating the tree-structured Parzen estimator TPESampler, was used for automated tuning of model hyperparameters over 100 trials. Data parsing and statistical calculations were conducted using pandas and numpy. Final model serialisation was performed via joblib.
  3. 2c. Computing infrastructure used: The computational pipeline, including hyperparameter search, nested cross-validation, and final model evaluation, was executed on the COSMOS cluster at LUNARC, Lund University's centre for scientific and technical computing. The experiments were run on standard CPU compute nodes with two AMD EPYC 7413 Milan processors operating at 2.65 GHz, providing 48 physical cores per node. System memory was 256 GB DDR4 RAM per node. Storage and interconnect used a high-speed HDR InfiniBand interconnect at 100 Gbit/node and a 3 PB IBM SpectrumScale parallel file system. Resource scheduling and job execution were managed through SLURM. The operating system was Rocky Linux 9 x86_64.
  4. 2d. README file for generating the results: See response to 2e.
  5. 2e. Reproduction script to produce all results reported in the paper: The process required training 20 predictors, for which feature selection was performed. Reproduction would require months of computing time. Some third-party scripts cannot be shared. Generated feature files can be provided upon request. Standard Python installations of the ML algorithms were used, and the parameters are listed in the Supplementary material.

6.3 Data quality

  1. 3a. Data sources and ground-truth annotations: The datasets, both for training and testing, were obtained from ClinVar and LOVD. There were two blind test sets. The review status for ClinVar cases was multiple submitters or higher. LOVD cases were asserted both by the submitter and the database curator.
  2. 3b. Set from which the dataset is sampled: The datasets were obtained from ClinVar and LOVD.
  3. 3c. Why the dataset is useful for the modeling task: ClinVar and LOVD are the largest high-quality variant databases available.
  4. 3d. Outcome variable and definition: PON-P4 classifies variants as pathogenic, benign, or VUSs. Pathogenic variants are considered disease-causing or disease-related; benign variants are not related to disease; and VUSs are variants of unknown significance with respect to pathogenicity. VUSs can be benign in some individuals and disease-related in others.
  5. 3e. Sample size and outcome frequencies: Training and test datasets:
    Details 10xCV set variations Blind test set variations Total variations Extended test set variations
    Pathogenic 11,405 1,246 12,651 1,743
    Benign 11,377 1,335 12,712 1,743
    VUS 11,421 1,291 12,712 1,743
    Total 34,203 variations; 3,588 proteins 3,872 variations; 374 proteins 38,075 variations; 3,962 proteins 5,229 variations; 851 proteins
    The blind test sets were disjoint from the training data. All variants in a protein were either in the training set or the test set.
  6. 3f. Percentage of missing data: 0%.
  7. 3g. Representativeness of the dataset: All pathogenic variants were collected, and a randomly selected matching number of benign cases and VUSs was included. Performance assessment on randomly selected blind test sets indicated the tools' generalisation ability.

6.4 Data preprocessing

  1. 4a. Excluded samples: Conflicting variants in ClinVar were excluded since they cannot be classified into a single class.
  2. 4b. Impossible or corrupt samples: A defined set of conditions was applied to obtain clinically classified cases.
  3. 4c. Transformations from raw dataset to model input: The variants were aligned with the MANE v1.3 reference sequences and annotated with the Variant Effect Predictor (VEP). Variants lacking cDNA, protein, or genomic information, as well as duplicates, were eliminated.

6.5 Modeling

  1. 5a. Models trained: The detailed description is in the Methods section and the main text of the article. Altogether, 20 predictors in six categories were trained and tested.
  2. 5b. Choice of model types implemented: Four algorithms were tested first: LightGBM, RF, SVM, and XGBoost. These algorithms were selected for their strong prior performance and algorithmic diversity. Since LightGBM performed best, it was used in all subsequent steps. An ESM2-based LLM was also trained for comparison.
  3. 5c. Model evaluation method: Initially, 10-fold cross-validation was used to evaluate the tools' performance. Two blind test sets were then used. For the train-test splits, see 3e.
  4. 5d. Model selection method: An analytical rubric with 10 measures was used. Seven of these were provided separately for the three variant types: B, P, and VUS. The method showing the overall best performance was chosen.
  5. 5e. Hyperparameter tuning: Hyperparameters of PON-P4 were optimised with Optuna 4.2.1 using the tree-structured Parzen estimator. Optimisation was conducted in two stages: hyperparameters were first tuned, and the final model was then trained on the best parameter set using 10-fold cross-validation to ensure robust generalisation.
  6. 5f. Appropriate baselines: The test sets were disjoint from the training data. The results for 10-fold cross-validation and the blind tests are practically identical.

6.6 Data leakage

  1. 6a. Training-only preprocessing and modeling: The blind test sets were used only after the methods had been trained on the training dataset.
  2. 6b. Dependencies or duplicates between training and test datasets: Duplicates were removed during data collection. Training was only done on training data. Features that showed high correlation were reduced to a single feature.
  3. 6c. Legitimacy of model features and leakage prevention: The features used describe either the variation, its context, or the affected gene or protein. These are valid characteristics for the application. There is no data leakage because the training and test datasets were disjoint, and the training data was not used to define or derive any features. Protein- and gene-level features describe properties of the affected gene or protein and were therefore identical for all variants within the same gene or protein. Variation-level features describe amino acid substitutions and their positions within the protein sequences. Structural features characterise variant positions within three-dimensional protein structures obtained from experimental structures or AlphaFold2 or AlphaFold3 models. Features with a standard deviation of 0 were removed, Pearson correlations were calculated for the remaining feature pairs, and redundant features were removed when correlation exceeded 0.8. The most significant features were then selected using RF-RFE until the highest prediction rates were attained.

6.7 Metrics and uncertainty

  1. 7a. Metrics used to assess and compare model performance: Performance assessment followed recommendations for variation interpretation tools and used a comprehensive set of performance measures: specificity, sensitivity, negative predictive value, positive predictive value, accuracy, Matthews correlation coefficient, overall performance measure, correct prediction ratio, generalised squared correlation, and overall accuracy. The first seven measures were split by variation type: B, P, and VUS.
  2. 7b. Uncertainty estimates and calculation: For the initial B and P prediction, PON-P3 was used. High-confidence predictions were identified with a probabilistic method using 200 parallel predictors. Since the probability distribution function of bootstrap probabilities is impossible to obtain, Chebyshev's inequality was used to calculate predictions within a certain standard deviation from the mean. The optimised LightGBM classifier, integrated with a random ECOC framework, was evaluated using 10-fold cross-validation on the 40-feature set. The Optuna-selected hyperparameters yielded a mean MCC of 0.6662 +/- 0.0101, ranging from 0.6400 to 0.6800, and a mean accuracy of 0.7756 +/- 0.0091. The relatively narrow standard deviations for MCC and accuracy indicated consistent predictive performance across training and validation partitions. Hyperparameter importance analysis indicated that tree depth, ECOC code size, and number of leaves were the principal determinants of model performance, together accounting for over 85% of the optimisation influence. The optimal ECOC configuration had a code_size of approximately 2.69, adding redundancy to improve robustness and reliability of the final multiclass predictions.
  3. 7c. Statistical tests and assumptions: The performance assessment measures were taken from widely used recommendations and cover the full spectrum of performance features.

6.8 Generalizability and limitations

  1. 8a. Evidence of external validity: The blind test sets were used to address external validity. PON-P4 has been generalised to predict variations in any human protein.
  2. 8b. Contexts where the findings may not hold: The method is intended for classifying amino acid substitutions in human proteins. If used for any other application, usability must be tested against relevant benchmarks.