Hi FDRBench developers,
This is a rather rudimentary question relative to the others on this forum as I am new to the tool.
I have used FDRBench to generate both* a shuffled (Human) and foreign (Human + A. thaliana) entrapment database, at the protein level. This was then used in Fragpipe using one of the in-built DIA workflows (DIA_Speclib_Quant). This workflow generates reports pre- and post-DIANN quant, but I am unsure which report is most appropriate for use in FDRBench.
*Edit to clarify: two FASTAs were generated for two independent searches
In order of increasing depth, some candidate outputs containing scoring information include:
- "protein.tsv": per protein reporting: Protein Probability, Top Peptide Probability, Protein Qvalue
- "peptide.tsv": per-peptide matrix reporting: Probability, Qvalue
- "ion.tsv": per-modified peptide matrix reporting: Probability, Qvalue, Expectation
- "psm.tsv" is a per-spectrum matrix reporting: SpectralSim, RTScore, Expectation, Hyperscore, Nextscore, Probability, Qvalue
- "report.tsv" (DIANN-generated) is a comprehensive per-run large matrix that reports many different values including: Q.Value, PEP, Global.Q.Value Lib.Q.Value, PG.Q.Value, PG.PEP, GG.Q.Value, Protein.Q.Value, Global.PG.Q.Value, Lib.PG.Q.Value
The last in the list is the only one generated by DIANN. As I have enabled MBR for DIANN within Fragpipe, I wonder whether that is the most appropriate. However, it is a rather large file (4.65GB). There is also report.parquet at 780 MB.
Can you please advise which output is appropriate, which column should be used as the basis for 'score', and what processing steps you'd recommend prior to use in FDRBench for FDP estimation?
Thank you for this tool and thank you in advance for any advice.
Anuk
Hi FDRBench developers,
This is a rather rudimentary question relative to the others on this forum as I am new to the tool.
I have used FDRBench to generate both* a shuffled (Human) and foreign (Human + A. thaliana) entrapment database, at the protein level. This was then used in Fragpipe using one of the in-built DIA workflows (DIA_Speclib_Quant). This workflow generates reports pre- and post-DIANN quant, but I am unsure which report is most appropriate for use in FDRBench.
*Edit to clarify: two FASTAs were generated for two independent searches
In order of increasing depth, some candidate outputs containing scoring information include:
The last in the list is the only one generated by DIANN. As I have enabled MBR for DIANN within Fragpipe, I wonder whether that is the most appropriate. However, it is a rather large file (4.65GB). There is also report.parquet at 780 MB.
Can you please advise which output is appropriate, which column should be used as the basis for 'score', and what processing steps you'd recommend prior to use in FDRBench for FDP estimation?
Thank you for this tool and thank you in advance for any advice.
Anuk