January 23, 2011

randomness


Microarray Home
Introduction
Services and Projects
GeneChip Expression
GeneChip Genotyping
Custom Microarrays
Data Analysis
People

Microarray Home

The Microarray Resource provides microarray analysis service and technical expertise to all researchers at Boston University and to interested groups from outside the university.  We are located on the 6th floor of the Evans (E) building at the BU Medical Campus. 

                                             

                   Affymetrix GeneChip                                 Custom Oligonucleotide Microarray

We offer two different microarray platforms at the Microarray Resource, Affymetrix GeneChips and custom Oligonucleotide microarrays.  For more information go to the Services [link] section of the webpage.  For both platforms users of the facility provide us with high quality samples and the Microarray Resource does the rest. 

Introduction

Services
      GeneChip Expression Analysis                       [Prices]
      GeneChip Genotyping Analysis                     [Prices]
      Custom Oligonucleotide Microarrays             [Prices]

Data Analysis

Microarray Resource People

Other information
Starting Material for Microarray Analysis
Protocols
Grant Materials
Links                     

All interested researchers are encouraged to contact the Microarray Resource to discuss a potential project or to ask any question about the use of microarrays.

Dr. Norman Gerry 617-414-1219                     npgerry@bu.edu
Dr. Marc Lenburg 617-414-1375                     mlenburg@bu.edu
Microarray Resource 617-414-1377

We look forward to talking with you!


Introduction

Biological arrays are an ordered set of compounds that are affixed to a solid surface.  By applying a sample solution to the array it is possible to assay the interaction of the sample with each of the compounds on the array.  Microarrays are a powerful research tool because they enable massively-parallel assays of biological samples.


The most common application of microarrays is gene expression analysis.  In this case the interaction between labeled mRNA in a biological sample and complementary DNA probes, which are affixed to the array, allows rapid quantification of the entire transcriptome.  Many other types of arrays are possible.  Nucleic acid arrays can also be used for genotyping and antibody arrays can be used for proteomics.

We offer two different microarray platforms at the Microarray Resource.  The Affymetrix GeneChip system is a commercial platform that enables rapid, reproducible, and accurate microarray analysis on a genome-wide scale.  The Affymetrix GeneChip platform can be used for both gene expression and genotyping analysis.

We also offer made-to-order microarrays that are manufactured by synthesizing oligonucleotide probes and spotting them onto glass microarray slides here at the Microarray Resource.  These “custom” microarrays are able to detect hundreds to thousands of genes of specific interest to individual investigators.  The custom microarrays are an extremely flexible platform.  They can be used for gene expression analysis of almost any organism in addition to many other types of projects.  Additionally, by making our own microarrays, we are able to significantly reduce the cost of performing microarray experiments.

At the Boston University Microarray Resource we strive to allow you to easily incorporate microarrays into your research and help you get the most out of your data.  You provide us with high quality samples and the Microarray Resource does the rest.


Services Offered by the Boston University Microarray Resource


Affymetrix GeneChips

The Affymetrix GeneChip™ system is a commercial microarray platform that allows whole genome gene expression analysis for common experimental organisms and high-throughput genotyping for human samples.

Custom Oligonucleotide Microarrays
The Microarray Resource will make custom arrays in-house by spotting oligonucleotide probes onto glass slides.  This is an extremely flexible platform allowing focused microarray analysis for any organism.


What types of projects can I use Microarrays for?

Affymetrix GeneChip Expression [link]
-          Genome wide expression analysis for common experimental organisms
o   Expression analysis to investigate cellular biology
o   Expression profiling to categorize biological samples

Custom Oligonucleotide Microarray [link]
-          Expression analysis for all organisms
o   Expression analysis to investigate cellular biology
o   Expression profiling to categorize biological samples
-          Chip on Chip
-          Gene Copy number
-          Spotting user provided cDNA, protein, anti-body, or small-molecule libraries

Affymetrix GeneChip Genotyping [link]
-          High Throughput Human Genotyping

Affymetrix GeneChip for Expression Analysis

The Affymetrix GeneChip system is a commercial microarray platform that allows whole genome gene expression analysis for common experimental organisms.  This system has three major advantages over other array systems.  It is easy to get rapid results, it has the capability to monitor the expression of every gene in the genome, and it is the most widely used commercial microarray platform.

However, the Affymetrix system also has a few disadvantages when compared with the Microarray Resource’s custom array system.  The GeneChip platform is significantly more expensive than custom microarrays, and Affymetrix only makes GeneChip arrays for common experimental organisms.

Getting started with Affymetrix GeneChips is easy
-           Set-up an appointment with members of the Microarray Resource to discuss your experiment.  This is optional, but highly recommended.
-           Give us 10 µg total RNA [more information about starting RNA link]
-           Wait 1 week for us to process your samples
-           Work with microarray core to analyze data

Available GeneChips for expression profiling and BUSM prices [Link] ]   Contact Us [Link]



Affymetrix GeneChip for Genotyping Analysis

Genotyping with Affymetrix GeneChips
The Affymetrix GeneChip™ system is a commercial microarray platform that allows high-throughput genotyping for human samples.  The GeneChip® Mapping 10K Array offers the ability to generate over 10,000 SNP genotypes from a single genomic DNA sample.  The 100K array, which will be released next year, will probe over 100,000 SNPs.



Getting started with Affymetrix GeneChips is easy
-           Set-up an appointment with members of the Microarray Resource to discuss your experiment.  This is optional, but highly recommended.
-           Give us 250 ng genomic DNA [more information about starting RNA link]
-           Wait 1 week for us to process your samples
-           Work with microarray core to analyze data

Please contact the microarray resource for more information about genotyping using the Affymetrix GeneChip platform

GeneChips for SNP genotyping and BUSM prices [Link]                Contact Us [Link]

Custom Oligonucleotide Microarrays

Our primary goal is to make microarray analysis more accessible to all researchers at BU.  We hope that microarray analysis of gene expression will become a method that researchers consider part of their regular repertoire of experimental approaches the way a Northern blot is now.  One of the most important ways that we can do this is by making the technology inexpensive.  Manufacturing custom Oligonucleotide microarrays in house will enabling an enormous cost savings.  The custom microarray system will allow researchers to choose a collection of genes of interest for their own research specific microarray.  Another advantage of custom microarrays is that while Affymetrix arrays are targeted for expression analysis and genotyping any sequence can be spotted on a custom array.  This opens up a wide range of additional application such as analysis of sequences enriched in chromatin immunoprecipitations, detecting region specific differences in copy number, pathogen detection, and many more.  Even in the area of gene expression analysis, custom microarrays have the advantage that they can be designed for any organism.


There are a number of ways to select genes for a custom microarray.  Known genes of interest can be the primary source but this approach can be supplemented with literature and database-mining or with preliminary experiments using Affymetrix whole-genome microarrays.

The Microarray Resource makes custom arrays in-house by spotting and covalently cross-linking oligonucleotide probes onto glass slides.  Probes for genes of interest are designed using sophisticated software that determines the best 50-70 nucleotide probe sequence for each gene.  These probes are then synthesized on an ABI 3900 high throughput DNA synthesizer.  The probes will be spotted onto derivatized glass slides using a Genetix QArray-Mini™ custom array spotter.  Once the arrays are made RNA samples from investigators are labeled and hybridized to these arrays and scanned with a Packard ScanArray Express™ multi-channel microarray scanner.  We will also spot user-provided libraries.

In order to make sure that you get the most out of your microarray experiment, the Microarray Resource will help you analyze your data.  This includes guidance on experimental design and statistical analysis, as well as access to software that will allow sophisticated data mining and visualization.

Getting started with Custom Arrays is easy
-           Set-up an appointment with members of the Microarray Resource to discuss your experiment.  This is optional, but highly recommended.
-           Determine a list of genes for your custom microarray
-           Wait a few weeks while we design and manufacture your microarrays.
-           Give us 5 µg total RNA [more information about starting RNA link]
-           Wait 1 week for us to process your samples
-           Work with microarray core to analyze data


Custom Oligonucleotide Microarray Prices [Link]                             Contact Us [Link]
Affymetrix GeneChip Prices for Expression Analysis

Affymetrix currently makes GeneChip expression arrays for 9 organisms.  The cost of the arrays varies from $300-$350 and is detailed below.  
Available Organisms
Human
Mouse
Rat

Yeast (cerevisiae)
Drosophila
P. aeruginosa
Arabidopsis
E. coli
C. elegans
B. subtilis
Barley

For human, mouse, and rat GeneChips, the entire transcriptome is split between two Affymetrix GeneChips.  In each case, the A chips contain the best annotated genes from the organism, while B chips contain mostly ESTs, splice variants, and poorly annotated transcripts.  The cost of running both A and B chips for human, mouse and rat samples is less than double the cost of running just the A chip because the processed RNA from a single sample can be hybridized to multiple arrays.

Organism
Genes
Annotated Genes
Chip
Reagents
Labor
Total
Human U133A
~22,500
19,993 w/ Gene Symbol
$350
$200
$300
$850
Human U133B
~22,500
10,043 w/ Gene Symbol
$350
$25
$150
$525
Mouse MOE430A
~22,500
?
$350
$200
$300
$850
Mouse MOE430B
~22,500
?
$350
$25
$150
$525
Rat ROE430A
~16,000
?
$350
$200
$300
$850
Rat ROE430AB
~16,000
?
$350
$25
$150
$525
Arabidopsis
~16,000
?
$300
$200
$300
$800
C. elegans
~22,500
?
$300
$200
$300
$800
Drosophila
~13,500
?
$300
$200
$300
$800
Yeast SG-98
~7,000
4,181 w/ Gene Symbol
$300
$200
$300
$800
E. coli
~5,500
?
$300
$200
$300
$800
P. aeruginosa
~6,000
?
$300
$200
$300
$800

Chip
Affymetrix GeneChips are available to us at the Boston Academic Consortium prices.  Due to the nature of our agreements with Affymetrix, we can only offer chips at these prices to academic investigators.  Other groups are encouraged to purchase their GeneChips from Affymetrix, and bring them to the microarray resource for hybridization.

Reagents and Labor
For each sample to be prepared for GeneChip analysis there is a cost of $500 for reagents and labor.  One sample is hybridized to each GeneChip microarray.  If the sample is to be hybridized to a set of A and B chips then the reagent and labor cost is $675.  Our charge for labor is very competitive with other facilities.  Remember, we provide comprehensive assistance with data analysis, a service unique to the Boston University Microarray Resource.

Experimental Design and Data Analysis
In order to make sure that you get the most out of your microarray experiment, the Microarray Resource will help you analyze your data.  This includes guidance on experimental design and statistical analysis, as well as access to software that will allow sophisticated data mining and visualization.

Affymetrix GeneChip Prices for Genotyping

Affymetrix currently makes one GeneChip for SNP genotyping.

Human Mapping 10K Array

Affymetrix anticipates releasing a similar 100K Mapping Array in 2004.


SNPs
Chip
Reagents & Labor
Total
Human 10K
>10,000
$400
?
?

Chip
Affymetrix GeneChips are available to us at the Boston Academic Consortium prices.  Due to the nature of our agreements with Affymetrix, we can only offer chips at these prices to academic investigators.  Other groups are encouraged to purchase their GeneChips from Affymetrix, and bring them to the microarray resource for hybridization.

Reagents and Labor
More information to come

Experimental Design and Data Analysis
In order to make sure that you get the most out of your microarray experiment, the Microarray Resource will help you analyze your data.  This includes guidance on experimental design and statistical analysis, as well as access to software that will allow sophisticated data mining and visualization.
Custom Microarray Prices

The cost of design and fabrication of custom microarrays is $150.  This cost includes probe selection, synthesis, and spotting. 

A minimum order is required on custom microarray projects, though you don't need to use -- or even necessarily print -- all of the arrays at one time.  The minimum order for a new custom array varies depending on the number of oligos in the array and the kind of array you want to make.
Minimum Order
Number of Oligos
Human, Mouse, Rat, and Yeast Expression Arrays
All Other Arrays
1-100
10 arrays
10 arrays
101-500
20 arrays
30 arrays
501-1000
40 arrays
60 arrays

Custom microarrays containing more than a thousand unique oligos are certainly possible. Investigators seeking to make arrays with more than a thousand unique oligos should contact us to discuss their project.  Discounts would be considered for projects larger than 100 arrays.  Investigators interested in projects of this size should contact us to discuss their project


Samples
Genes
Chip
Reagents
Labor
Total
Custom Microarray
2 per array
1-1,000 or more
$150
$100
$200
$450

Reagents and Labor
For each custom microarray there is a cost of $300 for reagents and labor.  Two samples are hybridized to each custom microarray.

Experimental Design and Data Analysis
In order to make sure that you get the most out of your microarray experiment, the Microarray Resource will help you analyze your data.  This includes guidance on experimental design and statistical analysis, as well as access to software that will allow sophisticated data mining and visualization.

Compared with the Affymetrix system, custom-array experiments can be done at greatly lower cost. The custom array itself costs only $150 per array, which is less than half of the cost of a GeneChip array.  Additionally, the cost of reagents and labor for hybridizing a custom array is $300 versus $500 for GeneChip arrays.  Another important cost savings with custom arrays is that two samples, such as control and experimental, are hybridized to a single array while just a one sample can be hybridized to a GeneChip array.  Consequently, the simplest custom array experiment costs $450 vs. $1600 for the simplest Affymetrix experiment.

Starting RNA for Affymetrix GeneChip Expression Analysis

-     10 µg high quality total RNA preferred.
-     Small amplification protocols are available that facilitate GeneChip expression analysis from samples as small as 100 ng.
-     In less than 10 µl of water (we will dry down dilute samples)
-     DNAse treatment is not necessary.  Small amounts of genomic DNA contamination will not affect the results of microarray analysis.
-     Poly-A selected RNA can be used for microarray analysis, but this is not recommended unless previous studies were conducted using poly-A selected RNA.
-     RNA extraction protocols [link]
-     The most common problem with RNA that we encounter in the microarray core facility is carryover organic contamination from the extraction.  This organic contamination will cause sample preparation reactions to fail. 


Starting DNA for Affymetrix GeneChip SNP Genotyping
-     250 ng high quality genomic DNA
-     DNA extraction protocols [link]

Starting RNA for Custom Microarray Analysis
-     5 µg high quality total RNA preferred.
-     In water
-     DNAse treatment is not necessary.  Small amounts of genomic DNA contamination will not affect the results of microarray analysis.
-     Poly-A selected RNA can be used for microarray analysis, but this is not recommended unless previous studies were conducted using poly-A selected RNA.
-     RNA extraction protocols [link]
-     The most common problem with RNA that we encounter in the microarray core facility is carryover organic contamination from the extraction.  This organic contamination will cause sample preparation reactions to fail. 

Microarray Protocols

GeneChip Expression Analysis Protocol [Link]
            Sample Preparation
            Hybridization, Staining, and Scanning
           
GeneChip Genotyping Protocol [Link]
            Sample Preparation
            Hybridization, Staining, and Scanning

Custom Oligonucleotide Microarray Protocol [Link]
            Array Production Protocols
            Sample Preparation and Hybridization Protocols

RNA extraction protocols

DNA extraction protocols

Data Analysis Methods



Data Analysis


Data Pre-processing

Introduction

Following Microarray hybridization and scanning there are a few things that need to be done to create a data set that is ready for analysis.  These include image quantification, normalization, and annotation.

Image Quantification

For Affymetrix GeneChips image quantification is performed using GeneChip Operating System 1.0 software (CGOS 1.0).  Starting with a scanned image GCOS determines the intensity of each 25mer probe on the GeneChip.  Then a gene specific intensity is calculated using the intensities of the set of probes for each gene.  This procedure is described in more detail in the Affymetrix Statistical Algorithms Reference Guide [https://www.affymetrix.com/support/technical/technotes/statistical_reference_guide.pdf]

Chip to Chip Normalization

Following within-chip image quantification, it is necessary to normalize the data across chips in order to make measurements as comparable as possible across chips.  There are a number of different normalization methods, but in general more complex methods will do a better job of normalization at the risk of overfitting.  Furthermore, as more samples are added to a microarray data-set, chip to chip differences become less important.  This makes complex normalization less important.  In general the Microarray Resource uses the simplest normalization method, linear scaling.

Linear Scaling

In linear scaling, the intensity of each gene on a chip is multiplied by a constant such that the average intensity of all the genes on that chip is scaled to a predetermined target.

Quantile Normalization

In quantile normalization, the intensity of each gene is ranked within each chip.  The average intensity across all chips of each rank is then calculated.  Finally, on each chip, the intensity of each gene is replaced by the average intensity of the gene of that rank across all chips.

Loess Normalization

In loess normalization, the intensity of the genes on a chip are normalized based on the local mean of signal intensities.
Gene to Gene Normalization
In addition to these normalization methods, which make chips comparable, there are other normalization techniques that make genes comparable on the same scale.  These methods are generally used prior to clustering or principle components analysis.

Log Ratio Normalization

In Log Ratio Normalization, the expression of each gene on each chip is calculated as;
log ( intensity of gene on this chip / mean intensity of gene across all chips)

Z-Score Normalization


Annotation

Introduction

In order to effectively analyze microarray data, it is critical for investigators to have access to complete and up-to-date annotation of the genes on the array.  At the Microarray Resource we get our annotation information from two primary sources, though there are a few others that are worth mentioning.

NetAffx

Affymetrix maintains the NetAffx [Link] database containing information about the genes that are contained on their GeneChip microarrays.  This is the best first source of information about Affmyetrix probe sets because each probe set has a unique page in the NetAffx database containing a broad range of information including gene and probe sequences, links to other databases, and functional descriptions of the genes.

Incyte Proteome Database

The Incyte Proteome BioKnowledge Library [Link] is now available for access by all current Boston University and Boston University Medical Center faculty, staff, and students.  This is an excellent database for finding information about genes from microarray experiments.  It is well curated and provides Pubmed links for all references.  This database is indexed by gene symbol.

Other Databases (NCBI etc.)

The are a number of other database that can provide valuable information about genes from microarray experiments
-                      Genbank
-                      SGI  (yeast)
-                      Gene Ontology

Identifying Differentially Expressed Genes

Introduction

With microarray data, biology researchers want to identify genes differentially expressed under different growth conditions or different treatments, to cluster genes according to their expression pattern, and to differentiate samples in pharmaceutical or clinical studies.

Fold Change

The most straightforward method of identifying differentially regulated genes in a microarray experiment is by fold change.  Fold change is the multiple by which the expression of a gene changed between two experimental groups.

Fold change can be reported using various scales that each convey the same information
Ratio: ¼, 4
Linear: -4, 4
Log base 2: -2, 2
Log base 10: -?, ?
Fold Change is usually calculated using the mean of a set of measurements within an experimental group, but I can also be calculated using the geometric mean, particularly if the original measurements were not converted to logarithmic scale.

While Fold Change is an important descriptor of the behavior of a genes expression between two experimental groups, it does not tell the whole story.  For example take the expression of one gene measured 4 times in each of two experimental groups.

Group A:         100, 200, 200, 300                  Mean = 200
Group B:         100, 100, 200, 2800                Mean = 800

Fold Change = 4

According to Fold Change this is a differentially regulated gene while we can see that Group B is not reproducibly upregulated 4 fold.  Consequently, Fold Change should not be used as a first pass method for identifying differentially expressed genes.

Statistical Significance

A better method for identifying differentially regulated genes is provided by statistics.  Analysis of Variance (ANOVA) is a technique that assesses whether a set of measurements from two or more experimental groups indicates, given observed variance, that the groups are different.  For microarrays the measurements are the expression levels of one gene and the groups correspond to the experimental sample groups.  ANOVA is used to identify genes that are differentially expressed in a manner that is reproducible across multiple measurements within each experimental group.

An ANOVA score is calculated by comparing the variance observed between the sample group means to the variance observed within the groups.  If the between group variance is high relative to the within group variance this indicates differential expression.  The result of an ANOVA is a probability, p, that an observed difference between groups could have been produced by chance if the groups were in fact the same.

Following the use of ANOVA to calculate a p-value for each gene it is useful to choose a p-value cut-off, below which genes will be considered differentially expressed, and above which genes will not be considered differentially expressed.  This cutoff will be arbitrary, but its’ choice should be made with an understanding of the trade-offs between sensitivity and selectivity that are inherent to choosing a significance cut-off.  In general, choosing a lower significance cut-off will result in fewer genes being identified as differentially expressed, but a smaller portion of those that are selected will be false-positives.  Choosing a higher significance cut-off will result in more genes being identified as differentially expressed, but a greater portion of those will be false-positives.  At any significance cut-off it is possible to estimate the associated false-positive and false-negative rates.  This allows an informed choice of the significance cut-off

ANOVA can take a few different forms depending on the experimental design.  The most basic type of ANOVA is a one-way ANOVA.  In a one-way ANOVA, the sample groups are stratified along a single experimental variable.  The simplest one-way ANOVA, with two sample groups, is equivalent to a T-Test.  The result of an ANOVA comparing more than two groups is the probability that any one of the groups is significantly different from the rest.  At the Microarray Resource we perform one-way as well as multiple-factor ANOVA.  Multiple-factor ANOVA differs from one-way ANOVA in that it generates p-value scores for each of the primary experimental axis as well as scores for each interaction between factors.

Multiple Hypothesis Testing

Correction of significance results for multiple hypothesis testing is an important concern in microarray data analysis.  It is common to use a p-value cut-off of 0.05.  In a microarray experiment in which 20,000 genes are measured, even if no genes are truly differentially expressed, 1,000 genes can be expected to meet the p < 0.05 significance cut-off by chance alone.  Furthermore, in the same 20,000 gene experiment with no changed genes, one unchanged gene would be expected to have a p-value as low as 0.00005.

A statistic test, like ANOVA, applied to microarray data tells you the probability that the observations made about a single gene could have been made if the null hypothesis, that the gene is not significantly changed, were true.  When applied to normally distributed random data, p-values will be evenly distributed between 0 and 1.  Thus, when looking at a single gene, a very low p-value is a significant finding, but as you increase the number of genes observed, the chance of finding a single very low p-value increases.

Take a fictitious microarray data set with 20,000 genes, none of which are differentially expressed between the experimental groups.  We will use a p-value cut-off of 0.05 to identify differentially regulated genes.  If we look at any one gene from our fictitious data set, which we know is not differentially expressed, there is a 1 in 20 chance of it having a p-value less than 0.05.  Our gene-wise false-positive rate, at this level of sensitivity, is 5%.  So, if we to use a microarray to observe the expression of a single gene, we can use p-value cut-off of 0.05 and control false positives at a rate of 5%.

If we use a statistical test and a p-value cut-off of p < 0.05 to identify differentially expressed genes from our fictitious microarray experiment, our gene-wise false positive rate is still 5%.  Five percent of 20,000 genes is 1,000 genes, that were not actually differentially expressed, but would be identified as significant at this level of sensitivity.  Testing as many hypotheses as there are genes on a microarray gives plenty of chances to make a mistake.

There are a few different methods for dealing with multiple hypothesis testing in significance analysis of microarray data.  The Bonferroni correction multiplies the significance observed for each hypothesis by the number of hypotheses being tested.  The Bonferroni correction is usually overly stringent for microarray data analysis.  If we use a Bonferoni corrected p-value cutoff of 0.05 on a real microarray data set, no matter how many genes meet the significance cut-off, there will be a 5% chance that a single false-positive will be among them.  If we identify 100 genes that are differentially expressed in an experiment, we would likely be willing to accept a few false-positives among the 100.  The Bonferonni criteria that there is only a 5% chance that a single false-positive is among the 100 is more control of false-positives than is usually necessary.  Increasing selectivity using the Bonferonni correction reduces sensitivity, so fewer differentially regulated genes will be identified.

Another method for treating the multiple hypothesis problem makes more sense for microarray experiments.  The False Discovery Rate (FDR) correction of Benjamini and Hochberg estimates the gene-wise false-positive rate among the genes at a significance cut-off.  The FDR is the quotient of the number of unchanged genes expected at a given significance cutoff over the number of genes detected at that significance cutoff.

The assumption that unchanged genes would have p-values evenly distributed between 0 and 1 can be used to estimate the number of false-positives expected at a given significance cut-off.  The number of false-positives expected at a given significance cut-off will be equal to the number of unchanged genes (or the number of genes on the microarray) times the p-value of the significance cut-off

Estimating the number of changed and unchanged genes

Based on two assumptions, it is possible to estimate the number of changed and unchanged genes in a microarray data set.  The first assumption is that unchanged genes will have p-values evenly distributed between 0 and 1.  The second assumption is that changed genes will not have p-values greater than a certain p-value threshold.

If there are no changed genes with p greater than the threshold then all of the genes with p greater than the threshold are unchanged.  If the unchanged genes have evenly distributed p-values, then the density of unchanged genes above the threshold will be the same as the density of unchanged genes below the threshold.  So, we calculate the density of unchanged genes above the threshold, and integrate this constant density from p equals 0 to 1.

Advanced Analysis Techniques

Principle Components Analysis

Technique

Principle Components Analysis is a mathematical transformation that can be applied to microarray data sets allowing data compression and dimensionality reduction.  The primary objective is to transform the data into a new space where data analysis is easier.  Princpal components analysis transforms a number of (possibly) correlated variables into a (smaller) number of uncorrelated variables called principal components.  The first principal component accounts for as much of the variability in the data as possible, and each succeeding component accounts for as much of the remaining variability as possible.
The mathematical technique used in PCA requires solving for the eigenvalues and eigenvectors of a microarray data-set in matrix form.  The eigenvector associated with the largest eigenvalue has the same direction as the first principal component. The eigenvector associated with the second largest eigenvalue determines the direction of the second principal component, etc..  The maximum number of eigenvectors equals the number of columns (samples) of the microarray data-set.
At the Microarray Resource we use principal components to view distributation of variability within the various samples that make up an experiment.
Figure Here


Looking at samples

Looking at genes

Hierarchical Clustering


Gene Clustering example (data before clustering / data after clustering)

Sample Clustering example (data before clustering / data after clustering)


At the Microarray Resource we use hierarchical clustering to visualize the expression profiles of a group of genes that have been selected using other statistical methods.

We perform hierarchical clustering using Spotfire software.  If you would like us to perform hierarchical clustering on your data-set, just give us a list of genes to cluster and we’ll do the rest.



K-Means Clustering

K-Means clustering is a technique that is used to divide genes into discrete groups


Biological Data Mining and Pathway Analysis

EASE

GenMapp and MappFinder

Visualization

Introduction

Visualizations are often associated with the presentation of microarray data.

Heat Map

  The most common of these visualizations is the heat map. 

Volcano Plot

In a Volcano Plot, the fold change and significance for each gene are displayed as a scatter plot.  Both fold change and significance are generally plotted in log scale.  The spots take a characteristic volcano form because absolute fold change is correlated with significance.

Volcano plots can be used to demonstrate fold change and significance cut-offs.
Picture here
Volcano plots are also an excellent way to visualize the changes that occur in a group of genes.
Picture here

Talk about making volcano plots comparing more than two groups?

Pathway Visualization

GenMapp


Other Crap

Oligo Design & Synthesis

The Microarray Resource will design oligonucleotide probes for detecting expression of specific genes of interest. This is not trivial as one must consider melting temperature, secondary structure, and sequence specificity, in addition to potential splice variants for each gene. We have automated many steps of this process.

The Microarray Resource will synthesize 50-70mers using the ABI 3900 DNA synthesizer at a rate of 100 oligos per day or more.  In contrast to cDNAs, which are commonly used as microarray probes, oligos provide flexibility to analyze the abundance of all mRNAs produced from a given gene.  One lesson from the large genome projects is that complexity may be generated, in part, by the surprisingly large number of mRNA splice variants derived from a single gene.  cDNA would not allow one to easily distinguish among different splice variants.

 Some pre-designed gene sets (e.g. 100 tumor suppressor genes) will soon be listed on the website to provide a starting point for figuring out what genes an investigator might want to analyze with their custom arrays.


Data Analysis
Data Warehousing & Data Analysis Consulting
Effectively managing and analyzing the large volume of data generated in each microarray experiment is a key factor to the successful use of the experimental approach.  One of our principal objectives as a microarray core facility is to provide software that will allow users of the Microarray Resource to get the most out of their data. 

C++
Netaffx
Excel
Spotfire

 

For Microarray Resource customers, ANOVA is implemented within Microsoft Excel.

Assumptions/Limitations of ANOVA

 

With microarray data, biology researchers want to identify genes differentially expressed under different growth conditions or different treatments, to cluster genes according to their expression pattern, and to differentiate samples in pharmaceutical or clinical studies.

GeneChip probe design

microarray data analysis methods


Data normalization



In order to make meaningful comparisons of a gene’s signal intensities between multiple microarrays it is necessary for the signal intensities from each chip to be expressed on the same scale. This is accomplished through data normalization.

The linear normalization provided by the Affymetrix software package MAS 5.0 serves to equilibrate the overall intensity levels of a group of chips, but it does not normalize intensity-dependent differences between chips.  Such differences are only dramatic on a small fraction or arrays and might result from the arrays being scanned such that some of the intensity values are outside the linear range of the detector.  Whatever the cause, we have found that a normalization scheme that can correct for intensity-dependent differences between chips results in a more accurate measure of signal intensity than that accomplished by linear normalization.  Consequently, we have employed a quantile method to normalize Affymetrix GeneChip microarrays.

The quantile normalization method adjusts the signal intensities on each chip as follows.  Within each array the signal intensity of each gene is ranked.  (example) If a set of genes is tied then each is given the mean of the set of ranks that the tied genes span.(example)  Across the arrays, the mean signal intensity for the genes of each rank is calculated.  For every gene on each array the signal intensity of the gene is replaced by the mean signal intensity for the genes of that rank across all of the arrays.

The effect of this normalization method is to make the distribution of signal intensities on each array identical.  The differences in the expression level of a particular gene between arrays are a result of where the measurement is ranked within each array.


Data filtering


The Affymetrix U133A and U133B arrays are capable of detecting the expression of a large fraction of the genes in the human genome.  As we expect that not every gene in the genome is expressed in endothelial cells, we sought to remove from our dataset both those genes that are not expressed in our samples as well as those that might be expressed but for which expression levels could not be reliably quantitated by the Affymetrix system.

To accomplish this task we took advantage of the fact that the probe set for each gene on the Affymetrix arrays contains eleven perfect match (PM) probes and an equal number of mismatch (MM) probes.  Unlike hybridization to the perfect match probes, hybridization to the mismatch probes is non-gene specific and the ratio of PM to MM hybridization indicates whether the PM hybridization is likely to be the result of gene-specific hybridization or is rather likely to be the result of experimental noise.  This concept has been mathematically formalized and algorithms for estimating the probability of PM hybridization resulting from gene-specific hybridization are provided in Affymetrix’s Microarray Suite 5.0 (MAS) software package.  Through the use of adjustable cut offs, MAS can further reduce these probabilities to a trinary “Present”, “Absent” or “Marginal” call – indicating high, low, or intermediate probability of gene-specific hybridization respectively.

We used the Affymetrix-recommended settings in MAS to call whether gene-specific hybridization was detected for each probe set on each array.  We then eliminated from our data set those probe sets for which gene-specific hybridization was “Absent” on every array.  These eliminated genes include both those that are not expressed in our endothelial cell line and those that cannot be detected due to technical limitations.  In our data set we included hybridization-intensity data from probe sets that had at least one “Present” call as we hypothesized that our different treatment conditions might cause expression of some genes to be induced from undetectable levels.

This filtering scheme eliminated xxxx genes from our data set (Table 1.)



Mass Spectrometry 793, Final Paper – Spring 2003


Garrett Frampton
Mass Spectrometry 793

Final Paper – Spring 2003



            It is usually the case that tissue samples subjected to proteomic or genomic analysis are comprised of many different cell types.  Frequently, the cells of interest are interspersed with other cells that the investigator does not want to examine.  These additional cells will certainly reduce the signal of molecules detected from the cells of interest and will likely confound analysis that are focused on only some of the cells in the tissue sample.  The ability to physically select cells of interest out of a tissue section is of great potential utility in many molecular biology applications and mass spectrometry is no exception.  Using laser capture microdissection this is possible.
In laser capture microdissection (LCM) cells can be selected individually out of a thin slice of frozen tissue that is less than one cell (5-20 mm) thick.  In LCM a thin ethylene vinyl acetate thermoplastic film is placed over the prepared tissue sample.  Using a microscope, the investigator targets the cells of interest with a laser.  The laser transiently heats and melts the film, adhering it to the cells below.  This can be repeated to capture as few a one and as many as a few thousand cells.  When the film is removed the captured cells remain adhered and are pulled out of the tissue section.  When performing LCM it is necessary to prepare the sample by actively removing all water in the tissue by treating it with organic solvents such as dehydrated ethanol or xylene.  Nevertheless, laser capture is a relatively gentle procedure, allowing recovery of many proteins and protein complexes, RNA transcripts several kilobases in length, and genomic DNA.
            The principle limitation associated with using laser capture microdissected cells for proteomic and genomic analysis is the small sample size.  Nature has provided a mechanism that allows amplification of single nucleic acid molecules, facilitating genomic analysis of microdissected samples.  Proteins, on the other hand, cannot be amplified.  Fortunately, mass spectrometry is an extremely sensitive detection method, and MS analysis LCM samples is currently viable.  Still, the small quantities of protein analyte that are obtained via LCM are a significant limit to proteomic analysis.
            In has usually been the case that proteins from tissue samples subjected to mass spectrometry are extracted into a lysis buffer.  The solubilized protein is then used for ESI or MALDI MS.  In addition to putting the protein in a state amenable to mass spectrometry, extraction of protein from tissue provides other benefits.  The dissolved protein can be subjected to purification and separation methods that will be discussed later.  Despite the benefits of extracting proteins from tissues, it does result in significant dilution.  Given the small amounts of protein that are procured from LCM samples, the dilution of that protein caused by extraction or separation could prevent detection of a protein of interest.
            Xu et al. describe a method by which cells obtained via LCM can be directly subjected to MALDI MS [1], obviating the need to extract proteins from the sample cells prior to mass spectrometry.  After capturing cells via LCM they attach the thermoplastic film to the target analysis plate and apply sinapinic acid matrix solution directly to the sample.  A portion of the cells’ protein associates with the matrix, making it amenable to ionization.  The sample is then directly subjected to MALDI-TOF mass spectrometry.
            This procedure has advantages to solubilizing protein prior to MS.  The first advantage is that the proteins are present at greater concentrations.  This translates directly to a higher signal to noise in mass spectra, which results in increased sensitivity, increased accuracy, and increased ability to detect proteins of interest.  Another related advantage is that with direct MALDI a very small original sample can be used for MS.  A third advantage is that direct MALDI MS of tissue samples is easier than extracting proteins prior to MS.  Eliminating the extraction creates a protocol that is less time consuming, less costly, less technically challenging, and less variable. 
            In fact, Xu et al. are not the first group to perform direct MALDI of LCM tissue.  Palmer-Troy et al. make an earlier report using a similar technique that they also use to investigate mammary carcinoma [2].  There is a difference between the two techniques with regard to the application of matrix solution to the MALDI target that Xu et al. claim gives their protocol an important advantage.  Palmer-Troy et al. apply 1 ml of matrix solution to the LCM cells on the target plate.  This volume far exceeds the volume of the cells.  In order to restrict matrix volume, Xu et al. use a finely pulled glass capillary to deposit matrix solution on LCM cells under microscopic visualization.  This allows them to employ matrix volumes of 100 pl to 10 nl.  Xu et al. claim that this gives their protocol an advantage because excessive matrix volume can dilute and mix the proteins associated with each cell cluster.  They also state that microspotting matrix solution provides an accurate target for the MALDI instrument’s camera system.  Xu et al. do not show a direct comparison of the two matrix deposition methods, but the spectra in their paper appear to be of higher quality than those of Palmer-Troy et al..  This could be attributed to the fact that Xu et al. are using a higher performance MALDI-TOF instrument though.
            In order to assess the quality of spectra obtained via their direct LCM MALDI MS protocol, Xu et al. perform comparisons between different methods of preparing the same tissue sample.  Their first figure shows that histological staining prior to LCM reduces spectral quality.  They show mass spectra of stained and unstained LCM captured cells.  The unstained cells show much better spectral quality.  This indicates that, if possible, LCM microscopy should be performed without staining.
The second figure states that the tissue dehydration process required in preparation for LCM does not result in significant protein loss from the tissues.  They show spectra obtained from three samples that were unwashed, washed in 70% ethanol for one minute, and washed in water for three minutes prior to direct MALDI MS without LCM.  The spectra appear similar, but this is not an adequate control.  The preparation of tissue for LCM involves not only a 70% ethanol wash, but also (2 x 30 sec) 95% ethanol washes, (2 x 30 sec) 100% ethanol washes, and a (1 x 5 min) xylene wash.  It is necessary to fully dehydrate the tissue in order to get good LCM results.  It seems likely that the stringent dehydration washes, which they do not perform in this control, could cause protein to be lost from the tissue.  Their statement, that the tissue dehydration process required in preparation for LCM does not result in significant protein loss from the tissue, is not tolerably supported.
            The third figure shows the mass spectrum obtained from a LCM sample of ten cells (mouse colon crypt) prepared using their protocol.  The spectrum looks good.  This indicates that their protocol will perform satisfactorily on as few as ten cells.
The fourth figure shows LCM capture of unstained human breast carcinoma.  The images show that almost all of the cancerous cells from the tissue section were removed without affecting the surrounding tissue.  This indicates that staining tissue samples is not required for accurate LCM.  This is important given their prior finding that staining reduces mass spectral quality.
            The fifth figure shows the mass spectra obtained by their direct MALDI MS protocol from microdissected invasive mammary carcinoma and normal breast epithelial cells.  This is a relevant comparison from a biological standpoint because epithelial cells are the primary cells from which breast cancer is derived.  They show spectra over a broad m/z range from 3,000 to 70,000 units.  The spectra appear generally similar with a few dozen clearly defined peaks.  Though the spectra are similar, there are clear differences between them.  There are a few peaks in one of the spectra that appear to be completely absent in the other.  Furthermore, many peaks common to both spectra are observed at significantly different intensities.  This indicates that there are differences between the protein content of these mammary carcinoma and normal epithelial samples, which is not surprising.
Comparison between different types, groups, or treatments of tissue is the basis of proteomic analysis via mass spectrometry.  These analyses generally fall into two categories.  The first type of analysis treats differences between the spectra being compared as markers of the samples being examined.  This can facilitate categorization of unknown samples.  In this case the identity and biological function of those proteins is not important.  The second type of comparative proteomic analysis is concerned with the identity and biology of the changed proteins.  Here the differences between the mass spectra of the samples, along with prior biological knowledge, are used as clues to help learn more about the underlying biology of the samples being examined.  Xu et al. do not comment on any analysis of the differences between the cancer and normal samples.  In this case the comparison between carcinoma and normal tissue provides evidence that their protocols can be used to perform comparative proteomics.
            The protocol described in Xu et al. can be used for rapid whole cell proteomics, but mass spectral analysis of whole cell protein extracts has an important fundamental limitation.  The mass spectra obtained from whole cell protein mixtures, even when acquired on state of the art mass spectrometers, is generally too noisy to detect most proteins of interest.  The most abundant cellular proteins, and their degradation products, obscure detection of other proteins, which are present in quantities many orders of magnitude lower than the abundant proteins.  Consequently, whole cell protein extracts are generally subjected to one or more separation or purification steps in combination with mass spectrometry.  The reduction in complexity by separation or purification can take many different forms.
            Tandem mass spectrometry is the most elegant method of separating proteins for mass spectrometry.  Quadrupole magnets can be used to trap a certain bandwidth of a mass spectrum, filtering a protein mixture by m/z.  In combination with CID, MSn is an excellent method to reduce mixture complexity for whole cell proteomics.
            Chromatography and electrophoresis are two other methods that are commonly used to separate protein mixtures.  They provide excellent results, but cause dilution of samples.  Chromatography and electrophoresis can also be used to separate protein mixtures across multiple dimensions, resulting in further reduction in mixture complexity.
            Affinity purification is another method that can be used to reduce the complexity of whole cell lysates.  In affinity purification, a favorable chemical interaction between an immobilized substrate and the protein(s) of interest is used to pull target molecules out of a complex mixture.  The form and selectivity of affinity capture methods is extremely variable.  Antibodies can be used to capture a particular protein species or a group of proteins.  Phosphorylated proteins can be purified using a variety of methods that convert phosphate groups to affinity handles.  Nucleic acids with sequence dependant specificity can be use to capture proteins.  Numerous inorganic compounds have also been demonstrated to have selective protein affinities.  Affinity purification is the basis of the SELDI technique that will be discussed later.
            Of these separation methods, only tandem mass spectrometry could be integrated into the direct MALDI protocol of Xu et al..  While tandem MS separates gas phase ions within the spectrometer, each of the other methods separates proteins in solution.  Since the protocol of Xu et al. skips the generation of lysate, none of the solution based separation methods would be possible.
            Craven et al. describe a protocol in which proteins are extracted into a lysis buffer and separated prior to mass spectrometry [3].  The protocol of Craven et al. is directly comparable to that of Xu et al. because both groups performed whole cell proteomic analysis of LCM samples, but Craven et al. used 2D-PAGE to simplify sample mixtures prior to MALDI.  They captured several thousand cells from normal kidney, renal carcinoma, and cervix epithelium and extracted proteins into a urea/thiourea-based lysis buffer.  They performed SDS poly-acrylamide gel electrophoresis and used silver staining to detect proteins within the gel.  Each spot was identified and excised by hand, protein subjected to in-gel tryptic digest, and fragment peptides extracted.  This resulted in a dilution of the whole cell protein extract into milliliters of solvent.  The extract from each gel spot was analyzed via MALDI-TOF mass spectrometry and peptide fragment masses were screened against the NCBI database.  Craven et al. identify between 470 and 930 gel features in their various samples, though they do not comment on how many proteins were actually identified.  Additionally, Craven et al. compare the results of using several different stains for LCM imaging, but they do not sample unstained tissue.  Given the evidence presented by Xu et al. that staining negatively affects spectral quality  it would have been good for Craven et al. to have included unstained tissue in this control.
            The 2D-PAGE protocol of Craven et al. generates far more spectral information than the direct MALDI protocol of Xu et al..  Craven et al. detect hundreds of proteins from a whole cell lysate and can obtain enough structural information to identify many of them.  This compares favorably to the protocol of Xu et al., which detects only a few dozen proteins and provides very limited structural information.
            While the 2D-PAGE method clearly has merits, the direct MALDI method also has strengths.  The most appealing feature of the direct MALDI method is it simplicity.  Assuming adequate ionization is achieved, the entire proteome could be ionized directly from tissue.  Compared to the involved chemistry of 2D-SDS-PAGE and MS of tryptic digests, direct MALDI is more likely to preserve proteins in their in vivo state.  Another extremely appealing characteristic of the direct MALDI method is that it allows more reasonable comparisons between proteins signals both within and across samples.  In the 2D-PAGE method each protein is subject to a separate excision, digest, extraction, and MS.  This makes it very difficult to control variability from these steps.  In the direct MALDI method the complexity of the mass spectra serves as an internal control.  Even if the direct MALDI method was naturally extended with MSn in order to detect more proteins, it would still be easier to reconstruct coherent proteome level information than it would be with 2D-PAGE and MS.  It is also important to realize that Craven et al. need to obtain a few thousand cells via LCM before they have enough protein to perform 2D-PAGE.  Xu et al., on the other hand, obtain quality mass spectra with as few as ten cells.  There is no application of LCM that would not benefit from having to obtain 10 cells instead of 10,000.  The tryptic peptide information provided by the 2D-PAGE method is currently invaluable for protein identification, but as knowledge of the proteome grows, identification of proteins without rich structural information will become easier.  As bioinformatics improves, 2D-PAGE and tryptic digest will no longer be necessary to identify most proteins.
            Surface enhanced laser desorption/ionization (SELDI) is another technique that allows reduction of mixture complexity for whole proteome analysis.  It requires relatively small quantities of protein compared with other separation and purification techniques, which makes it attractive for analysis of LCM samples.  In SELDI a complex protein mixture is applied to an affinity surface that subsequently serves as a target for MALDI.  It enables selective protein retention on a wide variety of surfaces.  Chromatographic surfaces include reversed phase, cationic, anionic and metal affinity while bioaffinity surfaces include antibody-antigen, DNA-protein, and receptor-ligand.  Following binding of target proteins to the affinity surface, an appropriate buffer is used to wash away non-specific binding and matrix is applied.  The bound proteins are then subjected to MALDI mass spectrometry.  SELDI is amenable to high throughput applications and can be configured in an array format to run many assays in parallel.  This makes it extremely attractive for many applications.
Batorfi et al. describe a protocol in which SELDI is used for whole cell proteome analysis from LCM tissue [4].  They capture a few thousand trophoblast cells from normal human placenta and hydatidiform mole (a malformed placental feature that generally results in termination of pregnancy) and extract whole protein into a modest 10 ml of lysis buffer.  3 ml of extract is applied to an immobilized metal affinity capture SELDI surfaces.  Following washing, sinapinic acid matrix is applied to the surface and MALDI-TOF MS is performed.  The IMAC surface binds proteins with exposed histidine, tryptophan, or cysteine residues.  This purifies a coherent fraction of the whole cell proteome and puts the bound proteins in a state amenable to mass spectrometry.  While the previous papers did not discuss the analysis of differentially detected ions, Batorfi et al. do find four peaks in their mass spectra that are significantly different between their two experimental groups.  They do not comment on the biological identity or relevance of these proteins.
Of these four papers, only one makes an attempt at the analysis of the biological significance of the data that they have collected.  This is an extremely important issue in current molecular biology.  Many new technologies allow collection of enormous amounts of data, which must be analyzed and deciphered before they are useful.  Unless technological boundaries are being pushed, the collection of proteomic and genomic data without equal investment into the analysis of that data is imprudent.  Considering the problematic nature of cross-experiment comparisons, it is incumbent upon researchers to conduct experiments with data analysis in mind.

Research Proposal


            Each of the methods discussed for conducting proteomic analysis of LCM tissue takes a different approach.  To me, the combination of tandem MS with direct MALDI is the most exciting because of its simplicity.  If I were to propose a research project for a graduate student to extend these findings, I would certainly investigate the application of LCM-MSn.
            LCM of cancerous tissue is particularly appropriate for two reasons.  First, cancerous tissue is a heterogeneous mix of diseased and normal tissue.  Analyzing the cancer cells separately might be necessary to figure out the molecular pathology of cancer.   Second, cancerous tissue is frequently available and diagnostic testing of cancer biopsies would both practical and valuable.  Consequently cancer tissue is a great starting point to investigate LCM-MSn.  Mammary carcinoma would be a good system because samples are available and the cell type of origin is known.
            Pushing the limits of LCM-MSn would require repeated analysis of the same samples.  To that end I would capture via LCM many dozen samples each of 100 cells from carcinoma and normal epithelial cells and freeze the samples at -80ºC for future use.  This could be problematic because long-term storage of LCM cells may result in protein degradation.  It would be important to monitor this factor and fresh LCM samples might have to be used.  I would also make sure that the original tissue was fresh and handled well in order to alleviate to problem of protein degradation prior to LCM.
            In order to perform tandem MS, it is necessary to define the desired bandwidth of m/z that will be captured in the first sector of the instrument.  There are two approaches to this problem.  One is to scan the entire range of m/z and subject each interval to a second stage of MS.  The second method is to use prior knowledge to select only those bands of m/z that contain important information.  I would investigate the scanning approach because it would be applicable to any type of sample.
The main problem with the scanning approach to tandem MS is that it requires collection of a separate CID mass spectrum at every bandwidth scanned.  This would result in thousands of spectra and hundreds of thousands of MALDI shots for a single 2D mass spectrum.  This would be difficult in terms of both sample size and time.  In order to efficiently investigate high-resolution tandem MS I would limit my investigation to a small portion of the initial mass spectrum.  I would choose a range of a few hundred m/z units, which contained a complex spectrum and possibly contained targets of interest that have been identified by other investigators.
The goal of this work would be to obtain a good 2D mass spectrum, over a limited range of initial m/z, using direct MALDI of LCM tissue.  The separation provided by tandem MS with CID would allow resolution of far more molecular species than in the report of Xu et al..  This would be an important extension of the direct MALDI ionization protocol because separation of complex protein extracts is critical for whole cell proteomics.
References

1.      Xu B, Caprioli R, Sanders M, Jensen R: Direct Analysis of Laser Capture Microdissected Cells by MALDI Mass Spectrometry. American Society of Mass Spectrometry 2002, 13:1292-1297

2.      Palmer-Troy DE, Sarracino D, Sgroi D, LeVangie R, Leopold P: Direct Acquisition of Matrix-assisted Laser Desorption/Ionization Time-of-Flight Mass Spectra from Laser Capture Microdissected Tissues. Clinical Chemistry 2000, 46:1513-1516

3.      Craven R, Totty N, Harnden P, Selby P, Banks R: Laser Capture Microdissection and Two-Dimensional Polyacrylamide Gel Electrophoresis.  American Journal of Pathology 2002, 160:815-822

4.      C Batorfi J, Ye B, Mok S, Cseh I, Berkowitz R, Fulop V: Protein profiling of complete mole and normal placenta using ProteinChip analysis on laser capture microdissected cells. Gynecologic Oncology 2003, 88:424-428

January 22, 2011

more random text


Embryonic stem cells or ES cells are cell lines derived from the inner cell mast of blastocyst stage embryos.  ES cells have two key properties that make them an ideal model system to study human development.  The first, self-renewal, is the ability to replicate almost indefinitely in cell culture.  This allows researchers to obtain large and homogenous populations of cells to study.  The second property, pluripotency, is the ability to differentiate into all of the of different cell types that make up an adult organism.

The science of regenerative medicine hopes one day to discover how to create healthy tissues to replace those that lose function due to disease, injury, or age. Since ES cells have the developmental potential to differentiate into any tissue are believed to hold great promise for regenerative medicine
In 2006 groundbreaking experiments by Takahashi and Yamanaka demonstrated that pluripotent cells could be created from somatic cells by forced expression of the ES cell transcription factors Oct4, Sox2, Klf4 and cMyc. These induced pluripotent stem cells, or iPS cells, are highly similar to embryonic stem cells.
. The generation of iPS cells was a huge breakthrough in regenerative medicine because they are similar to embryonic stem cells and can be derived in a patient-specific manner, from adult somatic cells, requiring no embryonic tissue. 
Since the initial reprogramming experiments in 2006, the pace of research in this field has been very rapid. In chimera experiments iPS were show to be able to give rise to germ line cells. Then in tetraploid complementation assays, it was demonstrated iPS cells can contribute to all of the tissues that make up an adult organism Since their initial generation in mouse, iPS cells have been made in several species including human .  They have been made from a large number of different starting cell types and using many different sets of ES cell transcription factors.
One of the original reprogramming transcription factors, Myc, is a known oncogene.  It has since been shown that iPS cells can be made without the use of myc. It has also been shown that iPS cells can be made with viral vectors than can later be excised from the genome or with vectors that do not integrate into the genome at all. iPS cells have also been made using small molecules and proteins as reprogramming agents.
Finally, a number of in vitro disease models have been made using iPS cells. For example, iPS cells were made from patients suffering from Parkinson's disease and then differentiated in vitro into neuronal cells. Since these neurononal cells have the genome of a person who developed Parkinson’s disease, they may be an important tool for understanding the disease pathology.
As I said earlier ES cells and iPS cells are highly similar. They express the same markers of pluripotency and have the same cell morphology. They also behave similarly in a broad range of phenotypic assays including            - embroid body formation           - teratoma formation - germline transmission             - and tetraploid complementation   Whether ES cells and iPS cell are truly equivalent is a very important question, since if they are equivalent, then the knowledge that has been gained in years of embryonic stem cell research will be applicable to iPS cells as well.
Recently, several studies have been published saying that there are differences in the gene expression programs of ES and iPS cells. If it is true, that ES and iPS cells have different gene expression programs, then this indicates that they are not equivalent cell types. This has huge implications for the entire field of stem cell research. If ES and iPS cells are not equivalent, then iPS cells might not hold the promise for regenerative medicine that we had hoped.
We decided that it would be critical to answer for ourselves whether ES and iPS cells are equivalent cell types.  We reasoned that obtaining genome-wide maps of H3K4me3 and H3K27me3 histone modifications in addition to gene expression profiles for a panel of ES and iPS cell lines would allow us to obtain a highly detailed and quantitative assessment of both the current transcriptional state of ES and iPS cells as well as their future developmental potential.
WE decided chose to profile the H3K4me3 and H3K27me3 histone modications because that are two of the most well studied histone modifications.  H3K4me3 is generally associated active genes while H3K27me3 is associated with genes that are not being transcribed.  In ES cells active genes have the K4 mark at their transcription start sites.  There is another class of genes that are occupied by both K4 and K27. These genes are especially interesting because they are include many of the most important genes for controlling differentiated cell types.
Homeobox transcription factors for example are nearly all occupied by both K4 and K27 in ES cells. These genes with both K4 and K27 are not expressed in ES cells, otherwise they would promote differentiation, but they are poised for later expression. In a differentiated cell that expresses one of these genes, the K27 mark is lost while K4 is retained.  In a differentiated cell that turns off one of these genes, the K4 mark will be lost and K27 is retained. At some genes, both K4 and K27 are retained which means that this gene is still poised for expression as the cell differentiates further.Together the genome wide locations of the K4 and K27 marks reflect both of the current transcriptional state of a cell as well as it future developmental potential.

We reasoned …
By comparing these data we can accurately asses whether there are truly differences between ES and iPS cells.  We assembled a collection of 6 independent ES cell lines each from separate donors, and 6 iPS lines, 4 derived from one fibroblast donor and 2 from another donor.   We also examined fibroblast cells as a control. 
Together H3K4me3 and H3K27me3 location analysis and microarray based gene expression experiments allow us to get a highly quantitative measure of cell state  This reflects both the current transcriptional state of the cell as well as its future developmental potential.  

January 20, 2011

MORE RANDOM GENOMICS TEXT


Methods for Comparison of ChIP-Seq experiments

Since the ChIP-Seq technique was first developed in 2007, a lot of work has gone into developing methods for analyzing this data.  Dozens of papers and several reviews have been published describing methods and tools for the analysis of individual ChIP-Seq experiments (references), but there has been far less effort devoted to methods for comparing the results from multiple ChIP-Seq location analysis experiments.  A comparison between multiple ChIP-Seq experiments might be necessary for examining the occupancy of a single factor in multiple conditions or for comparing the occupancy of different factors in the same cell.

Two of the methods that are commonly used to compared the results of multiple ChIP-Sqee experiments are Venn Diagrams and clsutergrams (Figure ?). Venn diagrams show the number of enriched regions or genes that overlap between a set of ChIP-Seq datasets. One weakness of Venn diagram analysis is that it requires determining a threshold for enrichment in each dataset a priori, which is often problematic. Additionally, the Venn diagram analysis frequently underrepresentats the similarity between two different experiments.  Two experiments that appear to be very similar might only be 50% overlapping in a Venn diagram analysis, and it is rare to observe two datasets overlapping more than 80%.  Lsutergrams show this and that etv.  The primary weakness of clsutergrams is tthey tend to highlight simialrities between datasets. 

Magnitude of enrichement more important than hypergeometricc significance
Framptongram - cytoscape
                        Methods for es v sips compairosn
                        Normalization and Comparitive Analysis

No threshold, preserve order