DDA and DIA identification using entrapment #373
Replies: 14 comments 22 replies
|
Dataset for testing: |
|
We should record the way of post-processing
|
|
Shouldn't be the level (PSM, peptidoform, protein) of the validation be on what was uploaded? For example, if you want to check for protein IDs, then the user should upload a table with protein accession names + their scores. Otherwise, ProteoBench carries out part of the data analysis, and thus would influence the outcome by whatever choices have been made. |
|
I put here the QQ-plot paper that was suggested in another discussion: https://pubs.acs.org/doi/10.1021/acs.jproteome.2c00423 (discussion https://github.com/orgs/Proteobench/discussions/205 now closed to be discussed here) |
|
We had a very productive discussion at the Lorentz Workshop on Trustworthiness in Proteomics on a ProteoBench module that would estimate random discovery proportion using entrapment. General idea:
Input data:
Metric
Participants must provide:
To define:
|
|
TODO: One thing that remains to be discussed/test:
|
|
Based on the discussion we had at the proteobench online meeting last week:
I will start setting up the module based on these decisions. If you object to any of them, let me know asap :) |
|
Update following a discussion in an online meeting (some new stuff, some that were already discussed here before):
|
|
We had a few meetings with interested contributors to refine the design of the module. Here are the meetings notes:
|
|
Follow up on our weekly ProteoBench meeting:
|
|
Hi @Cajac102, really great work!
|
|
we could remove it, the reason it exists is for cases where people upload FDRs that are not in the list and want to see the point anyways at that exact FDR. But that will probably never happen anyways. Ill remove it. |















Uh oh!
There was an error while loading. Please reload this page.
Since we decided to give up the entrapment strategy, we discussed a different type of DDA identification module:
It would be nice to plot overlap of identification from pairs of searches. It is too complicated to have it as a main figure: we would need to select pairs. So we think of a main figure to select from
What do we count? Ions? Peptidoforms? PSMs would be easier, but what do we do if we have multiple PSMs per spectrum: We keep the best one.
What file to use? Human sample?
All reactions