Changelog
All notable changes to ProteoPy will be documented in this file. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Changed
Plotting (
pr.pl):n_var_per_sample(), and with itn_peptides_per_sample()andn_proteins_per_sample(), now derive every axis and group order deterministically instead of following the order rows happen to occupy in the AnnData object. Bars, groups and blocks follow the category order of the annotation when the column is a Categorical, otherwise the lexicographic order of its values — so a user fixes the order once, by storing the annotation as an ordered Categorical, instead of relying on how the file was read. This changes existing figures: a plot whose bars followed.obs_nameswill now be sorted.x-axis labels come from
adata.obs["sample_id"]rather thanadata.obs_names. The drawn strings are unchanged —check_proteodata()enforces that the two are identical — but an AnnData axis index cannot carry a category order, so the column is the only source that can.ordernow subsets, as its documented semantics always claimed: values it omits are excluded from the plot and from the printed statistics, instead of being appended after the listed ones. Its values are validated againstadata.obs["sample_id"].ascendingcombined withorder_bysorts samples within each group; it was silently ignored before. It is still ignored, with a warning, whenorderorgroup_byis set.samples with a missing
order_byvalue are drawn in a trailing block labelledNAinstead ofnan.new
proteopy/pl/_utils.pyholdsresolve_default_order(), the shared implementation of the rule, for the remainingplfunctions to adopt.
Added
Preprocessing (
pr.pp):summarize_peptides_by_neighbourhood_union()collapses peptides that overlap in the protein sequence, keeping the most abundant member of each group. Peptide positions are resolved from a FASTA in the same call. A reimplementation of CCprofiler’ssummarizeAlternativePeptideSequences(topN = 1).groups by positional overlap rather than substring containment, so it sees peptide pairs that overlap without either containing the other, and needs no separate modification-summarisation step
selects the most abundant member instead of aggregating;
top_nsums the leading members instead, withkeep_lesscontrolling undersized groupsmissing values are deprioritised rather than removed: an incomplete peptide sorts last and loses to any complete competitor, but survives if its group has no complete member
equal totals are resolved by
tie_break_key, so the result does not depend on input row order; the default sorts non-letters after letters, favouring the unmodified form of an identifieron_unknown_proteinandon_unlocated_peptidedecide whether an unresolvable position raises, skips the peptide, or leaves the position undefined.varis reduced to the peptide-level proteodata columns plus the function’s own output, since the surviving row’s annotations describe one member rather than the group;keep_var_colscarries chosen columns through, aggregated across the group
Fixed
Datasets (
pr.datasets) and Download (pr.download):williams_2018()no longer conflates measured zeros with missing values. Three defects, each masked by another:zeros in
.Xwere coerced tonp.nan, discarding 13,547 genuine measurementscharge-state summation used pandas’ default
min_count=0, so a group with no measurements at all summed to0.0, inventing 3,324the same summation skipped
NaNinside a partially measured group, reporting a partial total as complete for 260 cellsthis changes
.X, and therefore the output ofpr.download.williams_2018(). Verified against the PRIDE deposit used by the original publication: the two are now bit-identical with an identical missingness pattern.
Changed
Datasets (
pr.datasets) and Download (pr.download):williams_2018()gained azero_to_naparameter (defaultFalse), for consistency with sibling functions. It is mutually exclusive withfill_na. Note it governs zero semantics only and does not restore the summation defects above.Preprocessing (
pr.pp):normalize_median()now defaults to
log_space=Truerenamed the
batch_idparameter togroup_byrenamed the
zeros_to_naparameter tozero_to_na, for consistency with sibling functionsno longer accepts sparse
.Xinput
[0.1.1] - 2025-03-24
Added
Preprocessing (
pr.pp):summarize_modifications()for modification summarizationAnalysis (
pr.tl): ANOVA support indifferential_abundance()Visualization (
pr.pl):binary_heatmap(),box(),volcano(),peptides_on_sequence(),peptides_on_prot_sequence();print_statsparameter across multiple plot functionsDatasets (
pr.datasets):williams_2018()andkarayel_2020()download functionsUtilities (
pr.utils): Public API withis_proteodata(),check_proteodata(),is_log_transformed()Documentation: Sphinx documentation site; proteoform inference and protein-level analysis tutorials
Changed
Reader (
pr.read):diann()now supports version >=1.9.1 with automatic version dispatchPreprocessing (
pr.pp):impute_downshift()now supportsgroup_by;normalize_median()gainsmethodparameter;remove_contaminants()defaults toinplace=TrueValidation:
is_proteodata()now checks for NaN in ID columns, infinite values in.X/layers, and obs/var index sync
Fixed
volcano_plottype incompatibility and label displayn_cat1_per_cat2_histminimum bin width
[0.1.0] - 2025-01-29
Initial release of ProteoPy.
Added
Data import (
pr.read): Support for DIA-NN and generic long-format tablesAnnotation (
pr.ann): Functions to annotate samples (.obs) and variables (.var)Quality control (
pr.pp): Completeness filtering, CV calculation, contaminant removalPreprocessing (
pr.pp): Median normalization, downshift imputationDifferential abundance (
pr.tl): t-test, Welch’s test, ANOVA with multiple testing correctionProteoform inference (
pr.tl): COPF algorithm reimplementation for detecting functional proteoform groupsVisualization (
pr.pl): Volcano plots, abundance rank plots, intensity distributions, CV plots, correlation matrices, hierarchical clustering profilesDatasets (
pr.datasets): Built-in example datasets (Karayel 2020)