Trim Reads
Sequencing reads are typically trimmed before use in downstream analyses such as mapping to a reference or de novo assembly. Trim Reads supports trimming based on a variety of factors, including:
- Quality trimming based on quality scores and removal of ambiguous nucleotides
- Adapter trimming using automatic read-through trimming and using an trim adapter list
- Homopolymer trimming for removal of homopolymers at the ends of reads
- Sequence trimming to remove specified numbers of bases at the 5' and/or 3' end or to trim reads to a fixed length
- Sequence filtering to remove reads based on a specified length threshold
The primary output from Trim Reads is a sequence list containing the trimmed reads. Optional outputs are a a sequence list containing discarded reads, a sequence list containing broken pairs, and a summary report. Further details about these outputs are provided in Trim output.
Trimming operation order
For each read supplied to Trim Reads, the effect of each trimming operation except for trimming to a fixed length, is calculated independently for the original read. The operation that would remove the most bases from the 5' end, and the operation that would remove the most bases from the 3' end, are then performed. This shortened read is then further trimmed to a fixed length, if that option was selected. Discarding reads based on a length threshold is done after all trimming operations have been performed.
Note that configuring trimming according to many criteria in a single run may occasionally expose an internal region in a read that could be further trimmed. In such cases Trim Reads may have to be run more than once, with different settings in each run.
An illustrative example
In this example, we consider the handling of a 100bp read, where the maximum length of reads after trimming has been set to 83 bp, with bases trimmed from the 3' end to achieve that length.
- First the effect of each trimming operation on the original read is calculated. In the list below, the highest numbers of bases to be trimmed from each end are in bold text for emphasis:
- Automatic read-through adapter trimming would remove 10 bases from the 5' end.
- Quality trimming would remove 8 bases from the 5' end.
- Ambiguous nucleotide trimming would removes the 3 bases from the 3' end.
- Homopolymer trimming would remove the 5 bases from the 5' end and 5 bases from the 3' end.
- Removing a fixed number of 5'/3' bases would remove 2 bases from the 5' end and 2 bases from 3' end.
- Trimming using the first adapter in a trim adapter list would remove 5 bases from the 5' end.
- Trimming using the second adapter in a trim adapter list would remove 2 bases from the 5' end.
- Based on the above, Trim Reads removes a total of 15 bases: 10 bases from the 5' end of the read and 5 bases from the 3' end. I.e. the maximum number of bases from each end, according to the calculations above.
- This initially trimmed read is 85 base pairs long. A fixed length of 83 was specified, so the read is further trimmed by removing the last 2 bases.
If at this point this read did not meet a specified length threshold, it would be discarded.
Launching Trim Reads
To start Trim Reads, go to:
Tools | Prepare Sequencing Data (
) | Trim Reads (
)
This opens a dialog where you can add sequences or sequence lists.
If multiple input elements are provided, each is processed separately.
The sections that follow describe the trimming options available and outputs that can be generated.
Subsections
- Quality trimming
- Adapter trimming
- Trim adapter list
- Homopolymer trimming
- Sequence trimming
- Sequence filtering
- Trim output
