Browse the manual

Introduction to CLC Cancer Research Workbench
- Contact information
- System requirements
  - Limitations on maximum number of cores
- Workbench Licenses
- About CLC Workbenches
  - New program feature request
  - Getting help
- When the program is installed: Getting started
  - Import of example data
- Plugins
- Network configuration
User interface
- View Area
- Zoom and selection in View Area
- Toolbox and Status Bar
- Workspace
- List of shortcuts
Data organization and management
- Navigation Area
- Customized attributes on data locations
- Filling in values
- Sequence web info
User preferences and settings
- General preferences
- Default view preferences
  - Number formatting in tables
  - Import and export Side Panel settings
- Data preferences
- Advanced preferences
  - Default data location
- Export/import of preferences
  - The different options for export and import
- View settings for the Side Panel
  - Saving, removing and applying saved settings
Printing
- Selecting which part of the view to print
- Page setup
  - Header and footer
- Print preview
Import/export of data and graphics
- Standard import
- Import tracks
- Import high-throughput sequencing data
- Import Primer Pairs
- Data export
- Export graphics to files
- Export graph data points to a file
- Copy/paste view output
History log
- Element history
  - Sharing data with history
Batching and result handling
- Batch processing
- How to handle results of analyses
  - Table outputs
  - Batch log
- Working with tables
  - Filtering tables
Viewing and editing sequences
- View sequence
- Circular DNA
  - Using split views to see details of the circular molecule
  - Mark molecule as circular and specify starting point
- Working with annotations
- Element information
- View as text
- Sequence Lists
Viewing structures
- Importing molecule structure files
- Viewing molecular structures in 3D
  - Moving and rotating
  - Troubleshooting 3D graphics errors
- Customizing the visualization
  - Visualization styles and colors
  - Project settings
- Snapshots of the molecule visualization
- Tools for linking sequence and structure
- Protein structure alignment
Getting started
- Reference data
- Create new folder
- Import data
  - How to import data
Preparing Raw Data
- Prepare sequencing data - all application types
- Analysis of sequencing data
Whole genome sequencing (WGS)
- Automatic analysis of sequencing data (WGS)
- Identify Variants (WGS)
  - How to run the 'Identify Variants' ready-to-use workflow
  - Output from the Identify Variants workflow
- Annotate Variants (WGS)
- Filter Somatic Variants (WGS)
- Identify Somatic Variants from Tumor Normal Pair (WGS)
- Identify Known Variants in One Sample (WGS)
Whole exome sequencing (WES)
- Automatic analysis of sequencing data (WES)
- Identify Variants (WES)
- Annotate Variants (WES)
- Filter Somatic Variants (WES)
- Identify Somatic Variants from Tumor Normal Pair (WES)
  - Import your targeted regions
  - How to run the 'Identify Somatic Variants from Tumor Normal Pair' ready-to-use workflow
- Identify Known Variants in One Sample (WES)
- Identify and Annotate Variants (WES)
Targeted amplicon sequencing (TAS)
- Automatic analysis of sequencing data (TAS)
- Identify Variants (TAS)
- Annotate Variants (TAS)
- Filter Somatic Variants (TAS)
- Identify Somatic Variants from Tumor Normal Pair (TAS)
  - Import your targeted regions
  - How to run the 'Identify Somatic Variants from Tumor Normal Pair' ready-to-use workflow
- Identify Known Variants in One Sample (TAS)
- Identify and Annotate Variants (TAS)
Whole Transcriptome Sequencing (WTS)
- Automatic analysis of RNA-seq data
- Analysis of multiple samples
- Annotate Variants (WTS)
- Compare variants in DNA and RNA
- Identify Candidate Variants and Genes from Tumor Normal Pair
- Identify variants and add expression values
- Identify and Annotate Differentially Expressed Genes and Pathways
Using data from other workbenches
- Open outputs from other workbenches
Genome browser tools
- Create new genome browser view
- Genome browser view
- Creating graph tracks
Quality control tools
- QC for Target Sequencing
- QC for Sequencing Reads
- QC for Read Mapping
  - Running the 'QC for Read Mapping' tool
  - Summary mapping report
Preparing raw data tools
- Merge overlapping pairs
  - Using quality scores when merging
  - Report of merged pairs
- Trim Sequences
- Demultiplex reads
Resequencing analysis tools
- Identify Known Mutations from Sample Mappings
  - Input and Parameters
  - Output from the 'Identify Known Mutations from Sample Mappings' tool
  - How to run the 'Identify Known Mutations from Sample Mappings' tool
- Trim primers of mapped reads
- Extract reads based on overlap
- Map Reads to Reference
  - Selecting reads and reference
  - Including or excluding regions (masking)
  - Mapping parameters
  - Gap placement
  - Computational requirements
  - Reference Caching
- Mapping output options
- Color space
  - Sequencing
  - Error modes
  - Mapping in color space
  - Viewing color space information
- Mapping result
  - View settings in the Side Panel
- Local realignment
  - Method
  - Realignment of unaligned ends
  - Guided Realignment
  - Multi-pass local realignment
  - Known Limitations
  - Computational Requirements
  - How to run the Local Realignment tool
- Merge mapping results
- Remove duplicate mapped reads
  - Algorithm details and parameters
  - Running the duplicate reads removal
- Coverage analysis
  - Running the Coverage analysis tool
- Variant Detectors - overview
  - Differences among the variants called by the three variant callers
  - How the variant detectors work
- Basic Variant Detection
- Fixed Ploidy Variant Detection
  - Ploidy and sensitivity
- Low Frequency Variant Detection
- Variant Detectors - error model estimation
- Variant Detectors - filters
  - General filters
  - Noise filters
- Variant Detectors - the outputs
  - The variant track output
  - The annotated table output
  - The report
- The Fixed Ploidy and Low Frequency variant callers: detailed descriptions
  - The Fixed Ploidy Variant caller: Models and methods
  - The Low Frequency Variant caller: Models and methods
- InDels and Structural Variants
  - How to run the InDels and Structural Variants tool
  - The Structural Variants and InDels output
  - The InDels and Structural Variants detection algorithm
  - The InDels and Structural Variants detection algorithm - Step 1: Creating Left- and Right breakpoint signatures
  - The InDels and Structural Variants detection algorithm - Step 2: Creating Structural variant signatures
  - Theoretically expected structural variant signatures
  - How sequence complexity is calculated
- Variant data
  - Variant tracks
  - The annotated variant table
  - Variant types
- Detailed information about overlapping paired reads
Add information to variants tools
- Add information from variant databases
- Add conservation scores
- Add exon number
- Add flanking sequence
- Add fold changes
- Add information about amino acid changes
- Add information from genomic regions
- Add information from overlapping genes
- Link Variants to 3D Protein Structure
- Download 3D Protein Structure Database
- From databases
Remove variants tools
- Remove variants found in external database
- Remove variants not found in external database
- Remove false positives
- Remove Germline Variants
- Remove reference variants
- Remove variants inside genome regions
- Remove variants outside genome regions
- Remove variants outside targeted regions
- From databases
Add information to genes tool
- Add information from overlapping variants
Compare samples tools
- Compare shared variants within a group of samples
- Identify Enriched Variants in Case vs Control Group
- Trio analysis
Identify candidate variants tools
- Create Filter Criteria
- Identify candidate variants
- Remove information from variants
- Identify variants with effect on splicing
Identify candidate genes tools
- Identify differentially expressed gene groups and pathways
- Identify highly mutated gene groups and pathways
- Identify mutated genes
- Select genes by name
Transcriptomics tools
- RNA-Seq analysis
- Small RNA analysis
- Experimental design
- Working with tracks and experiments
- Transformation and normalization
- Quality control
- Statistical analysis - identifying differential expression
- Feature clustering
  - Hierarchical clustering of features
  - K-means/medoids clustering
- Annotation tests
  - Hypergeometric tests on annotations
  - Gene set enrichment analysis
- General plots
Helper tools
- Extract sequences
Cloning and cutting
- Molecular cloning
- Gateway cloning
- Restriction site analysis
  - Dynamic restriction sites
  - Restriction site analysis from the Toolbox
- Gel electrophoresis
- Restriction enzyme lists
  - Create enzyme list
  - View and modify enzyme list
Sequencing Data Analysis
- Importing and viewing trace data
  - Scaling traces
  - Trace settings in the Side Panel
- Trim sequences
  - Trimming using the Trim tool
  - Manual trimming
- Assemble sequences
- Sort sequences by name
- Assemble sequences to reference
- Add sequences to an existing contig
- View and edit read mappings
- Reassemble contig
- Secondary peak calling
Primers
- Primer design - an introduction
  - General concept
  - Scoring primers
- Setting parameters for primers and probes
  - Primer Parameters
- Graphical display of primer information
  - Compact information mode
  - Detailed information mode
- Output from primer design
- Standard PCR
  - User input
  - Standard PCR output table
- Nested PCR
  - Nested PCR output table
- TaqMan
  - TaqMan output table
- Sequencing primers
  - Sequencing primers output table
- Alignment-based primer and probe design
- Analyze primer properties
- Find binding sites and create fragments
  - Binding parameters
  - Results - binding sites and fragments
- Order primers
Epigenomics
- ChIP-Seq Analysis
- Annotate with nearby gene information
Workflows
- Creating a workflow
- Distributing and installing workflows
- Executing a workflow
- Open copy of ready-to-use workflow
Legacy tools
- Quality-based variant detection
- Probabilistic variant detection
Appendix
- Use of multi-core computers
- Reference data overview
- Proteolytic cleavage enzymes
- Restriction enzymes database configuration
- Technical information about modifying Gateway cloning sites
- IUPAC codes for amino acids
- IUPAC codes for nucleotides
- Formats for import and export
  - List of bioinformatic data formats
  - List of graphics data formats
- SAM/BAM export format specification
  - Flags
- Gene expression annotation files and microarray data formats
- Translation Tables
- Matrices for alignment calculation
Bibliography

RNA-Seq analysis

Two tools are available for RNA-seq analysis, the tool RNA-Seq Analysis and the tool Create Fold Change Track Based on an annotated reference genome, the CLC Cancer Research Workbench supports RNA-Seq analysis by mapping next-generation sequencing reads and counting and distributing the reads across genes and transcripts. Subsequently, the results can be used for expression analysis using the tools in the Transcriptomics Analysis toolbox.

The tool for RNA-Seq analysis can be found here:

Toolbox | Transcriptomics Analysis () | RNA-Seq Analysis ()

The approach taken by the CLC Cancer Research Workbench is based on [Mortazavi et al., 2008].

The following describes the overall process of the RNA-Seq analysis when using an annotated eukaryote genome. See Specifying reads, reference genome and mapping settings for more information on other types of reference data.

The RNA-Seq analysis is done in several steps: First, all genes are extracted from the reference genome (using a gene track). Next, all annotated transcripts are extracted (using an mRNA track). If there are several annotated splice variants, they are all extracted.

An example is shown in figure 28.1.

Image rnaseq_explained1
Figure 28.1: A simple gene with three exons and two splice variants.

This is a simple gene with three exons and two splice variants. The transcripts are extracted as shown in figure 28.2.

Image rnaseq_explained2
Figure 28.2: All the exon-exon junctions are joined in the extracted transcript.

Next, the reads are mapped against all the transcripts plus the entire gene (see figure 28.3) and optionally to the whole genome.

Image rnaseq_explained3
Figure 28.3: The reference for mapping: all the exon-exon junctions and the gene.

From this mapping, the reads are categorized and assigned to the genes (elaborated later in this section), and expression values for each gene and each transcript are calculated.

Details on the process are elaborated in the following sections, which describe how to run RNA-seq analyses.

Subsections