Showing posts with label computational molecular biology. Show all posts

8.DRUG DESIGN



Drug design


Drug design, also sometimes referred to as rational drug design, is the inventive process of finding new medications based on the knowledge of the biological target.[1] The drug is most commonly an organic small molecule which activates or inhibits the function of a biomolecule such as a protein which in turn results in a therapeutic benefit to the patient. In the most basic sense, drug design involves design of small molecules that are complementary in shape and charge to the biomolecular target to which they interact and therefore will bind to it. Drug design frequently but not necessarily relies on computer modeling techniques.[2] This type of modeling is often referred to as computer-aided drug design.
The phrase '"drug design" is to some extent a misnomer. What is really meant by drug design is ligand design. Modeling techniques for prediction of binding affinity are reasonably successful. However there are many other properties such as bioavailability, metabolic half life, lack of side effects, etc. that first must be optimized before a ligand can become a safe and efficacious drug. These other characteristics are often difficult to optimize using rational drug design techniques.

Background

Typically a drug target is a key molecule involved in a particular metabolic or signaling pathway that is specific to a disease condition or pathology, or to the infectivity or survival of amicrobial pathogen. Some approaches attempt to inhibit the functioning of the pathway in the diseased state by causing a key molecule to stop functioning. Drugs may be designed that bind to the active region and inhibit this key molecule. Another approach may be to enhance the normal pathway by promoting specific molecules in the normal pathways that may have been affected in the diseased state. In addition, these drugs should also be designed in such a way as not to affect any other important "off-target" molecules or antitargets that may be similar in appearance to the target molecule, since drug interactions with off-target molecules may lead to undesirable side effects. Sequence homology is often used to identify such risks.
Most commonly, drugs are organic small molecules produced through chemical synthesis, but biopolymer-based drugs (also known as biologics) produced through biological processes are becoming increasingly more common. In addition mRNA based gene silencing technologies may have therapeutic applications.

[edit]Types

Flow charts of two strategies of structure-based drug design
There are two major types of drug design. The first is referred to as ligand-based drug designand the second, structure-based drug design.

[edit]Ligand based

Ligand-based drug design (or indirect drug design) relies on knowledge of other molecules that bind to the biological target of interest. These other molecules may be used to derive apharmacophore which defines the minimum necessary structural characteristics a molecule must possess in order to bind to the target.[3] In other words, a model of the biological target may be built based on the knowledge of what binds to it and this model in turn may be used to design new molecular entities that interact with the target.

[edit]Structure based

Structure-based drug design (or direct drug design) relies on knowledge of the three dimensional structure of the biological target obtained through methods such as x-ray crystallography or NMR spectroscopy.[4] If an experimental structure of a target is not available, it may be possible to create a homology model of the target based on the experimental structure of a related protein. Using the structure of the biological target, candidate drugs that are predicted to bind with high affinity and selectivity to the target may be designed using interactive graphics and the intuition of a medicinal chemist. Alternatively various automated computational procedures may be used to suggest new drug candidates.
As experimental methods such as X-ray crystallography and NMR develop, the amount of information concerning 3D structures of biomolecular targets has increased dramatically. In parallel, information about the structural dynamics and electronic properties about ligands has also increased. This has encouraged the rapid development of the structure-based drug design. Current methods for structure-based drug design can be divided roughly into two categories. The first category is about “finding” ligands for a given receptor, which is usually referred as database searching. In this case, a large number of potential ligand molecules are screened to find those fitting the binding pocket of the receptor. This method is usually referred as ligand-based drug design. The key advantage of database searching is that it saves synthetic effort to obtain new lead compounds. Another category of structure-based drug design methods is about “building” ligands, which is usually referred as receptor-based drug design. In this case, ligand molecules are built up within the constraints of the binding pocket by assembling small pieces in a stepwise manner. These pieces can be either individual atoms or molecular fragments. The key advantage of such a method is that novel structures, not contained in any database, can be suggested. These techniques are raising much excitement to the drug design community.[5][6][7]

[edit]Active site identification

Active site identification is the first step in this program. It analyzes the protein to find the binding pocket, derives key interaction sites within the binding pocket, and then prepares the necessary data for Ligand fragment link. The basic inputs for this step are the 3D structure of the protein and a pre-docked ligand in PDB format, as well as their atomic properties. Both ligand and protein atoms need to be classified and their atomic properties should be defined, basically, into four atomic types:
  • hydrophobic atom: all carbons in hydrocarbon chains or in aromatic groups.
  • H-bond donor: Oxygen and nitrogen atoms bonded to hydrogen atom(s).
  • H-bond acceptor: Oxygen and sp2 or sp hybridized nitrogen atoms with lone electron pair(s).
  • Polar atom: Oxygen and nitrogen atoms that are neither H-bond donor nor H-bond acceptor, sulfur, phosphorus, halogen, metal and carbon atoms bonded to hetero-atom(s).
The space inside the ligand binding region would be studied with virtual probe atoms of the four types above so the chemical environment of all spots in the ligand binding region can be known. Hence we are clear what kind of chemical fragments can be put into their corresponding spots in the ligand binding region of the receptor.

[edit]Ligand fragment link

Flow chart for structure based drug design
When we want to plant “seeds” into different regions defined by the previous section, we need a fragments database to choose fragments from. The term “fragment” is used here to describe the building blocks used in the construction process. The rationale of this algorithm lies in the fact that organic structures can be decomposed into basic chemical fragments. Although the diversity of organic structures is infinite, the number of basic fragments is rather limited.
Before the first fragment, i.e. the seed, is put into the binding pocket, and add other fragments one by one. we should think some problems. First, the possibility for the fragment combinations is huge. A small perturbation of the previous fragment conformation would cause great difference in the following construction process. At the same time, in order to find the lowest binding energy on the Potential energy surface (PES) between planted fragments and receptor pocket, the scoring function calculation would be done for every step of conformation change of the fragments derived from every type of possible fragments combination. Since this requires a large amount of computation, one may think using other possible strategies to let the program works more efficiently. When a ligand is inserted into the pocket site of a receptor, conformation favor for these groups on the ligand that can bind tightly with receptor should be taken priority. Therefore it allows us to put several seeds at the same time into the regions that have significant interactions with the seeds and adjust their favorite conformation first, and then connect those seeds into a continuous ligand in a manner that make the rest part of the ligand having the lowest energy. The conformations of the pre-placed seeds ensuring the binding affinity decide the manner that ligand would be grown. This strategy reduces calculation burden for the fragment construction efficiently. On the other hand, it reduces the possibility of the combination of fragments, which reduces the number of possible ligands that can be derived from the program. These two strategies above are well used in most structure-based drug design programs. They are described as “Grow” and “Link”. The two strategies are always combined in order to make the construction result more reliable.[5][6][8]

[edit]Scoring method

Structure-based drug design attempts to use the structure of proteins as a basis for designing new ligands by applying accepted principles of molecular recognition. The basic assumption underlying structure-based drug design is that a good ligand molecule should bind tightly to its target. Thus, one of the most important principles for designing or obtaining potential new ligands is to predict the binding affinity of a certain ligand to its target and use it as a criterion for selection.
Master Equation in Scoring Function.jpg
A breakthrough work was done by Böhm[9] to develop a general-purposed empirical function in order to describe the binding energy. The concept of the “Master Equation” was raised. The basic idea is that the overall binding free energy can be decomposed into independent components which are known to be important for the binding process. Each component reflects a certain kind of free energy alteration during the binding process between a ligand and its target receptor. The Master Equation is the linear combination of these components. According to Gibbs free energy equation, the relation between dissociation equilibrium constant, Kd and the components of free energy alternation was built.
The sub models of empirical functions differ due to the consideration of researchers. It has long been a scientific challenge to design the sub models. Depending on the modification of them, the empirical scoring function is improved and continuously consummated.[10][11][12]

[edit]Rational drug discovery

In contrast to traditional methods of drug discovery which rely on trial-and-error testing of chemical substances on cultured cells or animals, and matching the apparent effects to treatments, rational drug design begins with a hypothesis that modulation of a specific biological target may have therapeutic value. In order for a biomolecule to be selected as a drug target, two essential pieces of information are required. The first is evidence that modulation of the target will have therapeutic value. This knowledge may come from, for example, disease linkage studies that show an association between mutations in the biological target and certain disease states. The second is that the target is "drugable". This means that it is capable of binding to a small molecule and that its activity can be modulated by the small molecule.
Once a suitable target has been identified, the target is normally cloned and expressed. The expressed target is then used to establish a screening assay. In addition, the three-dimensional structure of the target may be determined.
The search for small molecules that bind to the target is begun by screening libraries of potential drug compounds. This may be done by using the screening assay (a "wet screen"). In addition, if the structure of the target is available, a virtual screen may be performed of candidate drugs. Ideally the candidate drug compounds should be "drug-like", that is they should possess properties that are predicted to lead to oral bioavailability, adequate chemical and metabolic stability, and minimal toxic effects. One way of estimating druglikeness is Lipinski's Rule of Five. Several methods for predicting drug metabolism have been proposed in the scientific literature, and a recent example is SPORCalc.[13] Due to the complexity of the drug design process, two terms of interest are still serendipity and bounded rationality. Those challenges are caused by the large chemical space describing potential new drugs without side-effects.

[edit]Computer-assisted drug design

Computer-assisted drug design uses computational chemistry to discover, enhance, or study drugs and related biologically active molecules. The most fundamental goal is to predict whether a given molecule will bind to a target and if so how strongly. Molecular mechanics or molecular dynamics are most often used to predict the conformation of the small molecule and to model conformational changes in the biological target that may occur when the small molecule binds to it. Semi-empirical, ab initio quantum chemistry methods, or density functional theory are often used to provide optimized parameters for the molecular mechanics calculations and also provide an estimate of the electronic properties (electrostatic potential, polarizability, etc.) of the drug candidate which will influence binding affinity.
Molecular mechanics methods may also be used to provide semi-quantitative prediction of the binding affinity. Alternatively knowledge based scoring function may be used to provide binding affinity estimates. These methods use linear regression, machine learning, neural nets or other statistical techniques to derive predictive binding affinity equations by fitting experimental affinities to computationally derived interaction energies between the small molecule and the target.
Ideally the computational method should be able to predict affinity before a compound is synthesized and hence in theory only one compound needs to be synthesized. The reality however is that present computational methods provide at best only qualitative accurate estimates of affinity. Therefore in practice it still takes several iterations of design, synthesis, and testing before an optimal molecule is discovered. On the other hand, computational methods have accelerated discovery by reducing the number of iterations required and in addition have often provided more novel small molecule structures.
Drug design with the help of computers may be used at any of the following stages of drug discovery:
  1. hit identification using virtual screening (structure- or ligand-based design)
  2. hit-to-lead optimization of affinity and selectivity (structure-based design, QSAR, etc.)
  3. lead optimization optimization of other pharmaceutical properties while maintaining affinity

Posted in | Leave a comment

7.DENDROGRAMS


DendrogramS


A dendrogram (from Greek dendron "tree", -gramma "drawing") is a tree diagram frequently used to illustrate the arrangement of the clusters produced by hierarchical clustering. Dendrograms are often used in computational biology to illustrate the clustering of genes.
For a clustering example, suppose this data is to be clustered using Euclidean distance as the distance metric.
Raw data
The hierarchical clustering dendrogram would be as such:
Traditional representation
Here the top row of nodes represent data, and the remaining nodes represent the clusters to which the data belong, and the arrows represent the distance.

Posted in | Leave a comment

3.MICRO ARRAYS


Microarray


A microarray is a multiplex lab-on-a-chip. It is a 2D array[clarification needed] on a solid substrate (usually a glass slide or silicon thin-film cell) that assays large amounts of biological material using high-throughput screening methods.
Types of microarrays include:
  • DNA microarrays, such as cDNA microarrays, oligonucleotide microarrays and SNP microarrays
  • MMChips, for surveillance of microRNA populations
  • Protein microarrays
  • Tissue microarrays
  • Cellular microarrays (also called transfection microarrays)
  • Chemical compound microarrays
  • Antibody microarrays
  • Carbohydrate arrays (glycoarrays)

DNA microarray

A DNA microarray is a multiplex technology used in molecular biology and in medicine. It consists of an arrayed series of thousands of microscopic spots of DNA oligonucleotides, called features, each containing picomoles (10−12 moles) of a specific DNA sequence, known as probes (or reporters). This can be a short section of a gene or other DNA element that are used to hybridize a cDNA or cRNA sample (called target) under high-stringency conditions. Probe-target hybridization is usually detected and quantified by detection of fluorophore-, silver-, or chemiluminescence-labeled targets to determine relative abundance of nucleic acid sequences in the target. Since an array can contain tens of thousands of probes, a microarray experiment can accomplish many genetic tests in parallel. Therefore arrays have dramatically accelerated many types of investigation.
In standard microarrays, the probes are attached via surface engineering to a solid surface by a covalent bond to a chemical matrix (via epoxy-silane, amino-silane, lysine, polyacrylamide or others). The solid surface can be glass or a silicon chip, in which case they are commonly known as gene chip or colloquially Affy chip when an Affymetrix chip is used. Other microarray platforms, such as Illumina, use microscopic beads, instead of the large solid support. DNA arrays are different from other types of microarray only in that they either measure DNA or use DNA as part of its detection system.
Example of an approximately 40,000 probe spotted oligo microarray with enlarged inset to show detail.

DNA microarrays can be used to measure changes in expression levels, to detect single nucleotide polymorphisms (SNPs) , to genotype or resequence mutant genomes (see uses and types section). Microarrays also differ in fabrication, workings, accuracy, efficiency, and cost (see fabrication section). Additional factors for microarray experiments are the experimental design and the methods of analyzing the data (see Bioinformatics section).

MicroRNA


MicroRNAs are a class of post-transcriptional regulators [1] [2] [3] . They are short ~22 nucleotide RNA sequences that bind to complementary sequences in the 3’ UTR of multiple target mRNAs, usually resulting in their silencing [4] . MicroRNAs target ~60% of all genes [5] , are abundantly present in all human cells [6] and are able to repress hundreds of targets each [7][8]. These features, coupled with their conservation in organisms ranging from the unicellular algae chlamydomonas reinhardtii [9] to mitochondria [10], suggest they are a vital part of genetic regulation with ancient origins [11].

MicroRNAs were first discovered in 1993 by Victor Ambros, Rosalind Lee and Rhonda Feinbaum during a study into development in the nematode C. elegans regarding the gene lin-14 [12]. This screen led to the discovery that the lin-14 was able to be regulated by a short RNA product from lin-4, a gene that transcribed a 61 nucleotide precursor that matured to a 22 nucleotide mature RNA which contained sequences partially complementary to multiple sequences in the 3’ UTR of the lin-14 mRNA. This complementarity was sufficient and necessary to inhibit the translation of lin-14 mRNA. Retrospectively, this was the first microRNA to be identified, though at the time Ambros et al speculated it to be a nematode idiosyncrasy. It was only in 2000 when let-7 was discovered to repress lin-41, lin-14, lin28, lin42 and daf12 mRNA during transition in developmental stages in c elegans and that this function was phylogenetically conserved in species beyond nematodes [13] [14], that it became apparent the short non-coding RNA identified in 1993 was part of a wider phenomenon. Since then over 4000 miRNAs have been discovered in all studied multicellular eukaryotes including mammals, fungi and plants. More than 700 miRNAs have so far been identified in humans [15] and over 800 more are predicted to exist. [16]. Comparing miRNAs between species can even be used to delineate molecular evolutionary history [17] on the basis that the complexity of an organism's phenotype may reflect that of the microRNA found in the genotype [18] .
The stem-loop secondary structure of a pre-microRNA from Brassica oleracea.
When the human genome project mapped its first chromosome in 1999, it was predicted it would contain over 100,000 protein coding genes. However, only around 20,000 were eventually identified (International Human Genome Sequencing Consortium, 2004) and for a long time much of the non-protein-coding DNA was considered "junk", though conventional wisdom holds that much if not most of the genome is functional [19]. Since then, the advent of sophisticated bioinformatics approaches combined with genome tiling studies examining the transcriptome[20], systematic sequencing of full length cDNA libraries [21] and experimental validation [22] (including the creation of miRNA derived antisense oligonucleotides called antagomirs) have revealed that many transcripts are for non protein coding RNA of which many new classes have been deducted such as snoRNA and miRNA [23] . Unfortunately, the rate of validation of microRNA targets is substantially more time consuming than that of predicting sequences and targets.
Due to their abundant presence and far-reaching potential, miRNAs have all sorts of functions in physiology, from cell differentiation, proliferation, apoptosis [24] to the endocrine system [25][26], haematopoiesis [27], fat metabolism [28], limb morphogenesis [29] . They display different expression profiles from tissue to tissue [30] , reflecting the diversity in cellular phenotypes and as such suggest a role in tissue differentiation and maintenance [31] [32].


Posted in | Leave a comment

2.GENOMICS


DNA sequence


A DNA sequence or genetic sequence is a succession of letters representing the primary structure of a real or hypothetical DNAmolecule or strand, with the capacity to carry information as described by the central dogma of molecular biology.
The possible letters are A, C, G, and T, representing the four nucleotide bases of a DNA strand — adenine, cytosine, guanine, thymine— covalently linked to a phosphodiester backbone. In the typical case, the sequences are printed abutting one another without gaps, as in the sequence AAAGTCTGAC, read left to right in the 5' to 3' direction. Short sequences of nucleotides are referred to asoligonucleotides and are used in a range of laboratory applications in molecular biology. With regard to biological function, a DNA sequence may be considered sense or antisense, and either coding or noncoding. DNA sequences can also contain "junk DNA."
Sequences can be derived from the biological raw material through a process called DNA sequencing.
In some special cases, letters besides A, T, C, and G are present in a sequence. These letters represent ambiguity. Of all the molecules sampled, there is more than one kind of nucleotide at that position. The rules of the International Union of Pure and Applied Chemistry (IUPAC) are as follows:[1]
Electropherogram printout from automated sequencer for determining part of a DNA sequence

  • A = adenine
  • C = cytosine
  • G = guanine
  • T = thymine
  • R = G A (purine)
  • Y = T C (pyrimidine)
  • K = G T (keto)
  • M = A C (amino)
  • S = G C (strong bonds)
  • W = A T (weak bonds)
  • B = G T C (all but A)
  • D = G A T (all but C)
  • H = A C T (all but G)
  • V = G C A (all but T)
  • N = A G C T (any)

gene idenfication:

4.1. Identification of Genes in a Genomic DNA Sequence

4.1.1. Prediction of protein-coding genes

Archaeal and bacterial genes typically comprise uninterrupted stretches of DNA between a start codon (usually ATG, but in a minority of genes, GTG, TTG, or CTG) and a stop codon (TAA, TGA, or TAG; alternative genetic codes of certain bacteria, such as mycoplasmas, have only two stop codons). Rare exceptions to this rule involve important but rare mechanisms, such as programmed frameshifts. There seem to be no strict limits on the length of the genes. Indeed, the gene rpmJ encoding the ribosomal protein L36 (Figure 2.1) is only 111 bp long in most bacteria, whereas the gene for B. subtilis polyketide synthase PksK is 13,343 bp long. In practice, mRNAs shorter than 30 codons are poorly translated, so protein-coding genes in prokaryotes are usually at least 100 bases in length. In prokaryotic genome-sequencing projects, open reading frames (ORFs) shorter than 100 bases are rarely taken into consideration, which does not seem to result in substantial underprediction. In contrast, in multicellular eukaryotes, most genes are interrupted by introns. The mean length of an exon is ~50 codons, but some exons are much shorter; many of the introns are extremely long, resulting in genes occupying up to several megabases of genomic DNA. This makes prediction of eukaryotic genes a far more complex (and still unsolved) problem than prediction of prokaryotic genes.

4.1.1.1. Prokaryotes

For most common purposes, a prokaryotic gene can be defined simply as the longest ORF for a given region of DNA. Translation of a DNA sequence in all six reading frames is a straightforward task, which can be performed on line using, for example, the Translate tool on the ExPASy server (http://www.expasy.org/tools/dna.html) or the ORF Finder at NCBI (http://www.ncbi.nlm.nih.gov/gorf/gorf.html).
Of course, this approach is oversimplified and may result in a certain number of incorrect gene predictions, although the error rate is rather low. Firstly, DNA sequencing errors may result in incorrectly assigned or missed start and/or stop codons, because of which a gene might be truncated, overextended, or missed altogether. Secondly, on rare occasions, among two overlapping ORFs (on the same or the opposite DNA strand), the shorter one might be the real gene. The existence of a long “shadow” ORF opposite a protein-coding sequence is more likely than in a random sequence because of the statistical properties of the coding regions. Indeed, consider the simple case where the first base in a codon is a purine and the third base is a pyrimidine (the RNY codon pattern). Obviously, the mirror frame in the complementary strand would follow the same pattern, resulting in a deficit of stop codons [235]. Figure 4.1 shows the ORFs of at least 100 bp located in a 10-kb fragment of the E. coli genome (from 3435250 to 3445250) that encodes potassium transport protein TrkA, mechanosensitive channel MscL, transcriptional regulator YhdM, RNA polymerase alpha subunit RpoA, preprotein translocase subunit SecY, and ribosomal proteins RplQ (L17), RpsD (S4), RpsK (S11), RpsM (S13), RpmJ (L36), RplO (L15), RpmD (L30), RpsE (S5), RplR (L18), RplF (L6), RpsH (S8), RpsN (S14), RplE (L5), and RplX (L24). Although the two ORFs in frame +1 (top line, on the right) are longer (207 aa and 185 aa) than the ORFs in frame −3 (bottom line, 117 aa, 177 aa, 130 aa, and 101 aa), it is the latter that encode real proteins, namely the ribosomal proteins RplR, RplF, RpsH, and RpsN.
Because of these complications, it is always desirable to have some additional evidence that a particular ORF actually encodes a protein. Such evidence comes along many different lines and can be obtained using various methods, e.g. the following ones:
  • The ORF in question encodes a protein that is similar to previously described ones (search the protein database for homologs of the given sequence).
  • The ORF has a typical GC content, codon frequency, or oligonucleotide composition (calculate the codon bias and/or other statistical features of the sequence, compare to those for known protein-coding genes from the same organism).
  • The ORF is preceded by a typical ribosome-binding site (search for a Shine-Dalgarno sequence in front of the predicted coding sequence).
  • The ORF is preceded by a typical promoter (if consensus promoter sequences for the given organism are known, check for the presence of a similar upstream region).

The most reliable of these approaches is a database search for homologs. In several useful tools, DNA translation is seamlessly bound to the database searches. In the ORF finder, for example, the user can submit the translated sequence for a BLASTP or TBLASTN (see 4.4) search against the NCBI sequence databases. In addition, there is an opportunity to compare the translated sequence to the COG database (see 3.4). A largely similar Analysis and Annotation Tool (http://genome.cs.mtu.edu/aat.html), developed by Xiaoqiu Huang at Michigan Tech [361], also compares the translated protein sequences to nr and SWISS-PROT; in addition, it checks them against two cDNA databases, the dbEST at the NCBI and Human Gene Index at TIGR.
Other methods take advantage of the statistical properties of the coding sequences. For organisms with highly biased GC content, for example, the third position in each codon has a highly biased (very high or very low) frequency of G and C. FramePlot, a program that exploits this skew for gene recognition [380], is available at the Japanese Institute of Infectious Diseases (http://www.nih.go.jp/~jun/cgi-bin/frameplot.pl) and at the TIGR web site (http://tigrblast.tigr.org/cmr-blast/GC_Skew.cgi). The most useful and popular gene prediction programs, such as GeneMark and Glimmer (see 3.1.2), build Markov models of the known coding regions for the given organism and then employ them to estimate the coding potential of uncharacterized ORFs.
Inferring genes based on the coding potential and on the similarity of the encoded protein sequences to those of other proteins represent the intrinsic and extrinsic approaches to gene prediction [110], which ideally should be combined. Two programs that implement such a combination, developed specifically for analysis of prokaryotic genomes, are ORPHEUS (http://pedant.gsf.de/orpheus [249]) and CRITICA ([67], source code at http://www.math.uwaterloo.ca/~jhbadger/). Several other algorithms that incorporate both these approaches are aimed primarily at eukaryotic genomes and are discussed further in this section.

4.1.1.2. Unicellular eukaryotes

Genomes of unicellular eukaryotes are extremely diverse in size, the proportion of the genome that is occupied by protein-encoding genes and the frequency of introns. Clearly, the smaller the intergenic regions and the fewer introns are there, the easier it is to identify genes. Fortunately, genomes of at least some simple eukaryotes are quite compact and contain very few introns. Thus, in yeast S. cerevisiae, at least 67% of the genome is protein-coding, and only 233 genes (less than 4% of the total) appear to have introns [660]. Although these include some biologically important and extensively studied genes, e.g. those for aminopeptidase APE2, ubiquitin-protein ligase UBC8, subunit 1 of the mitochondrial cytochrome oxidase COX1, and many ribosomal proteins, introns comprise less than 1% of the yeast genome. The tiny genome of the intracellular eukaryotic parasite Encephalitozoon cuniculi appears to contain introns in only 12 genes and is practically prokaryote-like in terms of the “wall-to-wall” gene arrangement [425]. Malaria parasite Plasmodium falciparum is a more complex case, with ~43% of the genes located on chromosome 2 containing one or more introns [272]. Protists with larger genomes often have fairly high intron density. In the slime mold Physarum polycephalum, for example, the average gene has 3.7 introns [851]. Given that the average exon size in this organism (165 ± 85 bp) is comparable to the length of an average intron (138 ± 103 bp), homology-based prediction of genes becomes increasingly complicated.
Because of this genome diversity, there is no single way to efficiently predict protein-coding genes in different unicellular eukaryotes. For some of them, such as yeast, gene prediction can be done by using more or less the same approaches that are routinely employed in prokaryotic genome analysis. For those with intron-rich genomes, the gene model has to include information on the intron splice sites, which can be gained from a comparison of the genomic sequence against a set of ESTs from the same organism. This necessitates creating a comprehensive library of ESTs that have to be sequenced in a separate project. Such dual EST/genomic sequencing projects are currently under way for several unicellular eukaryotes (see Appendix 2).

4.1.1.3. Multicellular eukaryotes

Figure 4.2
Figure 4.2
Organization of the human iduronate 2-sulfatase gene(more...)
In most multicellular eukaryotes, gene organization is so complex that gene identification poses a major problem. Indeed, eukaryotic genes are often separated by large intergenic regions, and the genes themselves contain numerous introns, many of them long. Figure 4.2 shows a typical distribution of exons and introns in a human gene, the X chromosome-located gene encoding iduronate 2-sulfatase (IDS_HUMAN), a lysosomal enzyme responsible for removing sulfate groups from heparan sulfate and dermatan sulfate. Mutations causing iduronate sulfatase deficiency result in the lysosomal accumulation of these glycosaminoglycans, clinically known as Hunter's syndrome or type II mucopolysaccharidosis (OMIM entry 309900) [896]. A number of clinical cases have been shown to result from aberrant alternative splicing of this gene's mRNA, which emphasizes the importance of reliable prediction of gene structure [631].
Obviously, the coding regions compose only a minor portion of the gene. In this case, positions of the exons could be unequivocally determined by mapping the cDNA sequence (i.e. iduronate sulfatase mRNA) back to the chromosomal DNA. Because of the clinical phenotype of the mutations in the iduronate sulfatase gene, we already know the “correct” mRNA sequence and can identify various alternatively spliced variants as mutations. However, for many, perhaps the majority of the human genes, multiple alternative forms are part of the regular expression pattern [118,576], and correct gene prediction ideally should identify all of these forms, which immensely complicates the task.
Ideally, gene prediction should identify all exons and introns, including those in the 5′-untranslated region (5′-UTR) and the 3′-UTR of the mRNA, in order to precisely reconstruct the predominant mRNA species. For practical purposes, however, it is useful to assemble at least the coding exons correctly because this allows one to deduce the protein sequence.
Correct identification of the exon boundaries relies on the recognition of the splice sites, which is facilitated by the fact that the great majority of splice sites conform to consensus sequences that include two nearly invariant dinucleotides at the ends of each intron, a GT at the 5′ end and an AG at the 3′ end. Non-canonical splice signals are rare and come in several variants [329,582]. In the 5′ splice sites, the GC dinucleotide is sometimes found instead of GT. The second class of exceptions to the splice site consensus includes so-called “AT-AC” introns that have the highly conserved /(A,G)TATCCT(C,T) sequence at their 5′ sites. There are additional variants of non-canonical splice signals, which further complicate prediction of the gene structure.
The available assessments of the quality of eukaryotic gene prediction achieved by different programs show a rather gloomy picture of numerous errors in exon/intron recognition. Even the best tools correctly predict only ~40% of the genes [697]. The most serious errors come from genes with long introns, which may be predicted as intragenic sequences, resulting in erroneous gene fission, and pairs of genes with short intergenic regions, which may be predicted as introns, resulting in false gene fusion. Nevertheless, most of the popular gene prediction programs discussed in the next section show reasonable performance in predicting the coding regions in the sense that, even if a small exon is missed or overpredicted, the majority of exons are identified correctly.
Another important parameter that can affect ORF prediction is the fraction of sequencing errors in the analyzed sequence. Indeed, including frameshift corrections was found to substantially improve the overall quality of gene prediction [133]. Several algorithms were described that could detect frameshift errors based on the statistical properties of coding sequences [224]. On the other hand, error correction techniques should be used with caution because eukaryotic genomes contain numerous pseudogenes, and non-critical frameshift correction runs the risk of wrongly “rescuing” pseudogenes. The problem of discriminating between pseudogenes and frameshift errors is actually quite complex and will likely be solved only through whole-genome alignments of different species or, in certain cases, by direct experimentation, e.g., expression of the gene(s) in question.


SNP,s and applications


In molecular biology and bioinformatics, a SNP array is a type of DNA microarray which is used to detect polymorphisms within a population. A single nucleotide polymorphism (SNP), a variation at a single site in DNA, is the most frequent type of variation in the genome. For example, there are around 10 million SNPs that have been identified in the human genome[1]. As SNPs are highly conserved throughout evolution and within a population, the map of SNPs serves as an excellent genotypic marker for research.

Principles


The basic principles of SNP array are the same as the DNA microarray. These are the convergence of DNA hybridization, fluorescence microscopy, and solid surface DNA capture. The three mandatory components of the SNP arrays are:
  1. The array that contains immobilized nucleic acid sequences or target;
  2. One or more labeled Allele specific oligonucleotide (ASO) probes;
  3. A detection system that records and interprets the hybridization signal.
To achieve relative concentration independence and minimal cross-hybridization, raw sequences and SNPs of multiple databases are scanned to design the probes. Each SNP on the array is interrogated with different probes. Depending on the purpose of experiments, the amount of SNPs present on an array is considered.

Applications


An SNP array is a useful tool to study the whole genome. The most important application of SNP array is in determining disease susceptibility and consequently, in pharmacogenomics by measuring the efficacy of drug therapies specifically for the individual. As each individual has many single nucleotide polymorphisms that together create a unique DNA sequence, SNP-based genetic linkage analysis could be performed to map disease loci, and hence determine disease susceptibility genes for an individual. The combination of SNP maps and high density SNP array allows the use of SNPs as the markers for Mendelian diseases with complex traits efficiently. For example, whole-genome genetic linkage analysis shows significant linkage for many diseases such as rheumatoid arthritis, prostate cancer, and neonatal diabetes. As a result, drugs can be personally designed to efficiently act on a group of individuals who share a common allele - or even a single individual. A SNP array can also be used to generate a virtual karyotype using specialized software to determine the copy number of each SNP on the array and then align the SNPs in chromosomal order.
In addition, SNP array can be used for studying the Loss of heterozygosity (LOH). LOH is a form of allelic imbalance that can result from the complete loss of an allele or from an increase in copy number of one allele relative to the other. While other chip-based methods (e.g. Comparative genomic hybridization) can detect only genomic gains or deletions, SNP array has the additional advantage of detecting copy number neutral LOH due to uniparental disomy (UPD). In UPD, one allele or whole chromosome from one parent are missing leading to reduplication of the other parental allele (uni-parental = from one parent, disomy = duplicated). In a disease setting this occurrence may be pathologic when the wildtype allelle (e.g. from the mother) is missing and instead two copies of the mutant allelle (e.g. from the father) are present. Using high density SNP array to detect LOH allows identification of pattern of allelic imbalance with potential prognostic and diagnostic utilities. This usage of SNP array has a huge potential in cancer diagnostics as LOH is a prominent characteristic of most human cancers. Recent studies based on the SNP array technology have shown that not only solid tumors (e.g. gastric cancer, liver cancer etc) but also hematologic malignancies (ALL, MDS, CML etc) have a high rate of LOH due to genomic deletions or UPD and genomic gains. The results of these studies may help to gain insights into mechanisms of these diseases and to create targeted drugs.

Posted in | Leave a comment