Bioinformatics training
Systems biology and omics technology
B A S I C L E V E L
In the last decades the rapid progress of the molecular biology and the achievements in this scientific area is directly related to the advances in the computer sciences and the development of new software tools.
Systems Biology and Omics Technologies: The Big Picture
In the last decades the rapid progress of the molecular biology and the achievements in this scientific area is directly related to the advances in the computer sciences and the development of new software tools. The application of different omics-related technologies (genomics, epigenomics, transcriptomic, proteomics, metabolomics, etc.) leads to the accumulation of versatile biological data and their combined interpretation open new possibilities for the scientists to move from studying isolated biological molecules towards a broad analysis of large sets of biological molecules. The biological sciences have become big-data sciences which are of growing importance for solving problems concerning human health and environment. The big data challenges are not only their size but also their increasing complexity.
Systems biology of the cell: from single-omics to multi-omics experiments data
Systems biology has emerged as a new interdisciplinary study field which has gained a tremendous attention in the last few years. Although the term “Systems biology” has been used in various ways, it is generally understood to describe research that combines biology with seemingly disparate disciplines such as physics, biochemistry, engineering, biostatistics, mathematics, computer science, bioinformatics and others. The main purpose of this cross-disciplinary field is to obtain a highly parallel view of the biological systems at the molecular level and understand how they function as a whole instead as a sum of parts. By clarifying the molecular mechanisms and processes at different system levels in the organism or a cell, it could be predicted how these systems change over time and under varying conditions. The great amount of information obtained on molecular bases of cell physiology and organization could provide solutions concerning diseases, toxicities, therapies, drug discovery, etc.
Systems biology has launched as a distinct discipline in 1950s with the work of the systems theorists Mihajlo Mesarovic and Ludwig von Bertalanffy on the general systems theory. The main concept of this theory is that the dynamics of any system is a result of the relationships between its separate units which determining its function. At that time, by thoroughly studying enzymes and kinetics of enzyme reactions, biochemists, followers of the system theory, tried to examine the behavior of the biochemical pathways as a network instead as a sum of its parts. Moving away from the biochemical area, Reinhart Heinrich developed theoretical approaches for the description and quantitative investigation of signaling pathways and developed a metabolic control theory. The need for an integrated approach for a more detailed study and understanding of the complex biological processes became evident.
Systems biology is highly dependent on the biological information obtained by molecular biology and / or individual omics approaches, but often they are largely hypothesis-driven or reductionist. In the reductionist approach the functional properties of the individual molecular components in the complex biological systems are studied. After the completion of the Human Genome Project the modern science has developed beyond the gene-centered view of the earlier genomic era. This period designated as a postgenomic era is characterized by major changes in the scientific research conduction and results interpretation and according to the Bloom this is the end of the “naïve reductionism”. Nevertheless, that these classical methods alone cannot provide a complete understanding of living organisms, they will continue to be an essential element of all biological research. In this context, Systems biology has caused the fundamental change in traditional approaches. It uses a new hypothesis-generating holistic approach, focuses on the study of subsystems, which allows deeper understanding of the whole process. The galactose utilization (GAL) pathway of yeast Saccharomyces cerevisiae is one example of the application of Systems biology approach for rereading and interpretation of the results obtained by the reductionist approach (one gene/one protein at a time). On the basis of experimental data obtained from protein and RNA levels analyses as well as protein–protein and protein–DNA interactions and their integration into a single model, new hypothesis was provided on the regulation of GAL pathway, which was experimentally verified afterward.
Through the use of computational and mathematical tools a large amount of experimental data is collected and integrated by Systems biology for revealing unknown patterns and hypothesis generation. This is an important strategy to gain new insights of the biological systems and also help in the experimental design.
The systems-level approach of the biological systems is aimed at clarifying the following three issues in the context of molecular network: i) what are the individual components of the system; ii) how do they work separately? and iii) how do these components work together to accomplish a task? (3). In the context of molecular networks, the basic purpose of the Systems biology can be summarized as follows: (i) an understanding of the structure of all the components of a cell / organism up to molecular level, (ii) the ability to predict the future state of the cell / organism under a normal environment, iii) the ability to predict the output responses for a given input stimulus, and iv) the ability to estimate the changes in system behavior upon perturbation of the components or the environment.
The term “systems” in Systems biology defines different span of complexity ranging from two macromolecules that interact to perform a particular task to whole organisms (Fig.1).

Thus, Systems biology shares a common scientific goal with the discipline of physiology which is dedicated to study the integrated function of the entire complex biological systems. Organisms are much more than the sum of their parts and the complexity of physiological processes cannot be understood simply by studying how the components work in isolation. For example, genes encode the primary structure of the cellular proteins that in turn carry out specific functions supporting cell metabolism and physiology and organism development. At the same time, the proteins in each cell do not act in isolation but in a complex network whose are important to specify phenotypes. Most biological processes are highly dynamic and often involve more than one type of molecules. At the same time one phenotype may be conditioned by several different molecular and epigenetic mechanisms and one type of molecule may be involved in different phenotypes. Actually, even in the same organism, a protein may play different roles in different cells; or signal pathway effectors may induce various differentiation programs in different cell lineages. Moreover, in multicellular organisms, single cells do not have an existence independent from the whole organism, they are ontogenetically linked.
Regardless of their common goal, Systems biology and physiology utilize different tools and experimental approaches, which leads to obtaining a different set of experimental data. A cross-disciplinary Systems biology is a modern rapidly evolving discipline due to the fact that it uses a variety of methods and tools including large-scale functional genomics and other omics technologies, bioinformatics and computer modeling, which are not exploited by the physiologists. The role of computational biology in Systems biology is to process and analyze massive amounts of empirical data produced by different omics levels which in turn leads to biological knowledge discovery and generation of research hypothesis. The application of in silico and simulation-based analyses give opportunities to make predictions which further are confirmed by experimental assays. The rapid advance of information technologies including also enhancement of public web-databases of biological information also support the development of omics studies and respectively lead to the progress of systems biology.
In order to comprehensively study biological processes, it is critical to understand how the separate biological layers (genome, epigenome, transcriptome, proteome, metabolome, and ionome) interconnect with one another in a cellular system and how the flow of biological information takes place. The combination of different omics analyses employing multi-omics approach is required to design a precise picture of living organisms. Transcriptomics, proteomics, and metabolomics data can answer key biological questions regarding the expression of transcripts, proteins, and metabolites, independently, but a systematic multi-omics integration can comprehensively assimilate, annotate, and model these large data sets.
Different branches of omics – challenges to combine biological information
The transfer of genetic information in biological systems is realized from DNA to mRNA to protein and this is stated as the Central dogma in molecular biology. The three main processes in each cell are replication, transcription and translation. Their constant flow ensures the maintenance and conversion of the genetic information, encoded in DNA into gene products, which are either RNAs or proteins, depending on the gene. Replication is a process of a cell`s DNA duplication and it is the basis for biological inheritance. It is carried out by the enzyme DNA polymerase that copies a single parental double-stranded DNA molecule into two daughter double-stranded DNA molecules. The enzyme RNA polymerase creates an RNA molecule from DNA and that process is known as transcription. The newly synthesized RNA molecule is complementary to a gene-encoding stretch of DNA. Translation makes protein from mRNA. The ribosome generates a polypeptide chain of amino acids using mRNA as a template. The polypeptide chain folds up to become a protein. In eukaryotic cells, or those cells that have a nucleus, replication and transcription take place within the nucleus while translation takes place outside of the nucleus in cytoplasm. In prokaryotic cells, or those cells that do not have a nucleus, all three processes occur in the cytoplasm. The organisms` phenotype is determined by that information transfer paradigm. Biologists have studied these “omes” for years in the form of genomics, transcriptomics and proteomics. The data from these experimental approaches are complemented by epigenomics and metabolomics that have recently been used to solve specific problems concerning many functions of an organism. The rapid development and advance in “omics” technologies determine the progressive expand of the volume of information that can be gathered in individual studies. Moreover, the current high throughput nature of these techniques has increased accessibility to this information in terms of time and cost. Many researchers are placed in a situation where they can collect several omics data sets on the same experimental samples. In aim to obtain more comprehensive conclusions on biological processes these data sets must be integrated by multi-omics approach and analyzed as a holistic system (Fig. 2).

The term “Omics” derived from a Greek word and the addition of thе suffix -ome to cellular molecules, such as gene, transcript, protein, metabolite, gives meaning to “whole,” “all,” or “complete.” The different layers of the cell, consisting of DNA and modifications (Genome, Epigenome), RNA and protein content (Transcriptome, Proteome), small molecules (Metabolome, Lipidome) and elemental composition (measured as ‘Ionome’), can be analysed by omic technologies. The combination of omic layers in a multi-ome dataset are integrated through robust systems biology which is able to reveal inter-layer mechanisms and interactions as well as the function of cell populations’ tissues, organs, and the whole organism. Omics approaches comprise a larger number of measurements per endpoint and although the number of parameters measured per analysis is increased, the number of replicates is decreased. This on one hand is due to the super estimation of methods since it is considered that more measurements would compensate a small number of samples and on the other hand is due to the cost and time of omics experiments.
The single-omic disciplines aimed at studying specific biological issues without requiring a prior understanding of the biological bases involved. Depending on the type of biomolecule they primarily focus on in a specific biological sample, omics technologies are divided into: genomics (genome / gene), metagenomics (genomes recovered directly from environmental samples), epigenomics (supporting structure of genome, including protein and RNA binders, alternative DNA structures, and chemical modifications on DNA), transcriptomics (mRNA), proteomics (peptides / proteins), metabolomics (metabolites), lipidomics (lipids) glycomes (carbohydrate and sugars), ionomics (ions).
Although none of the current omic technologies is perfect, some of them succeed to provide more comprehensive picture of the biological layer they aim to study than others. This fact is due not only to differences in the state of technological developments, but rather of the differences in the in chemical and physical complexity of each biological level.
Genomics
Genomics is the systematic study of all of an organism’s genes (the genome), including interactions of those genes with each other and with the organism`s environment. The genome is the basal biological layer in the cell and represents the total DNA of a cell or organism. Deoxyribonucleic acid (DNA) is the chemical compound that contains all the necessary genetic information needed to develop and direct the activities of nearly all living organisms. DNA molecules are made of two twisting, paired strands, often referred to as a double helix. Each DNA strand is made of four nucleotide bases – adenine (A), thymine (T), guanine (G), and cytosine (C) that pair specifically on opposite strands: A always pairs with a T; a C always pairs with a G. A sequence of three adjacent nucleotides (codon) encodes for a specific amino acid during protein synthesis or translation. The nucleotides` order along the DNA molecules determines the genetic code. The genetic code is universal because it is the same among all organisms and it is also degenerate because 64 codons encode only 22 amino acids. With its four-letter language, DNA contains the information needed to build an entire organism. A gene represents the unit of DNA that carries codes for making a specific protein or a set of proteins. On the basis of the intrinsic complementarity of the nucleotide bases in DNA, it has become possible to rapidly sequencing a huge number of genomes at a relatively low cost. From the efficiently sequenced genomes predictions about RNA and protein sequences could be made which is combined in the multi-omic approaches. The digital form of the DNA sequences as a kind of biological omic information could be easily stored in biological databases and shared between scientists all over the world. The sequencing of the first whole genome of the bacterium Haemophilus influenza in 1995 made a revolution in molecular biology. The big volume of sequences data was produced which was beyond being completely interpreted. To decipher relevant genetic information from the background genetic material was almost an impossible task. To overcome this, more detailed biological information was required. Information on the transcription of the genetic material and subsequent production of proteins was necessary.
An essential research direction in the field of Systems biology is the functional genomics. This discipline develops and exploit large-scale and high-throughput methodologies in aim to define and analyze gene function at a global level. It is very important for developing a systems level understanding of a biological process to identify the genes, and the proteins they encode which work together to give rise to that process. Functional genomics is an integrative scientific field which combines multiple large-scale datasets in attempts to generate insights into gene function. At first, genes were analyzed individually, but the advance of technologies in recent years make possible the expression of thousands of genes to be analyzed simultaneously. This large-scale analysis of gene function is called DNA microarrays technology (Fig. 3).

DNA microarrays measure differences in DNA sequence between individuals and this technology can provide information for the function of uncharacterized genes and also can reveal clusters of interacting genes that give rise to a biological process of interest. Microarray analyses can also provide insights into mechanisms of gene regulation, evolution and the etiology of disease. For example, the microarray data analysis could reveal abnormalities such as chromosomal insertions and deletions or abnormal chromosomal numbers in a process called comparative genomic hybridization. The most common variations in DNA sequences between people are single nucleotide polymorphisms (SNPs), in which one nucleotide is substituted for another; this may have functional significance if the change results in a codon for a different amino acid. They are of particular interest when linked with diseases with a genetic determination. Single nucleotide polymorphism profiling also has a role in pharmacogenomics in exploring individual patient responses to drugs.
Epigenomics
The epigenetic modifications such as DNA methylation, histone modifications, 2D and 3D analysis of chromatin structure and non-coding RNA are studied by the methods of epigenomics. This -omics direction focuses on the analysis of overall epigenetic changes which provides important information regarding mechanisms and function of gene regulation across many genes in a cell or organism. Scientists have understood that the individual`s phenotype is not controlled only by the genome but also by the changes in regulation of gene activities. Genetic experiments in humans and animals have proved that, in addition to the DNA sequence, epigenetic marks may be transmitted from parent to offspring via the gametes and influence the phenotype of offspring.
Epigenomics defines the modifications in the regulation of gene activities that act without, or independently of, changes in gene sequences. Some definitions confine epigenetics/epigenomics to modifications of the phenotype without changes of the DNA sequence that are transmitted to the next generations. The epigenetic modifications are chemical modifications, which are not coded by the genome and they coordinate how and when genes are expressed. The epigenomics explores heritable, reversible modifications of DNA and chromatin that do not influence the primary nucleotide sequences. While the term epigenomics would describe the analysis of epigenetic changes across many genes in a cell or throughout an entire organism, epigenetics centers on processes that regulate how and when specific genes are turned on and turned off. Several factors are known to affect the epigenetic regulation: 1) Nutrition (dietary factors); 2) Environmental factors; 3) Radiation exposure; 4) Infectious agents; 5) Immunological factors; 6) Genetic factors; 7) Toxic agents; 8) Mutagens.
Versatile detection methods are applied for analysis of epigenetics modifications in the cell. DNA methylation are analyzed by digestion assays and bisulfite sequencing of DNA. In digestion assays, the genomic DNA is fragmented with methylation-sensitive and methylation-insensitive endonucleases. Methylation-sensitive restriction enzymes cleave only unmethylated DNA and leave methylated DNA fragments undigested. The fragmented DNA can be analyzed by sequencing or microarray, and thus the methylation sites are mapped. The disadvantage of this method is that it only studies the DNA sequences near the targets of chosen restriction enzymes, and is usually used to characterize only global levels of DNA methylation instead of identifying methylated DNA at the single residue level. Resolution is improved by using the method of bisulfite sequencing. This strand-specific method is used to convert unmethylated cytosine to uracil, whereas methylated cytosine residues remain unaffected. The resulting DNA is amplified in PCR and can be analyzed by sequencing of the regions of interest. The analysis can also be done using MALDI-TOF mass spectrometry or microarrays. DNA methylation can also be studied through methods based on the principle of affinity chromatography. The sample containing fragmented DNA is loaded to a column with bound methyl-binding domain (MBD) of MeCP2, specific for methylated DNA. The methylated DNA fractions are eluted out and analyzed with genome-wide techniques such as MBDCap-seq/MethylCap-seq. Another approach is immunoprecipitation of methylated DNA (MeDIP), which is based on the specific binding of antibodies to the methylated cytosine (5mC) in DNA. This method is also used for the analysis of hydroxylated methylcytosine (5hmC). Purified fragments are analyzed by PCR, sequencing or microarray.
A widely used technique for detection of histone modifications is the immunoprecipitation of chromatin (ChIP). It is used for identifying local posttranslational modifications of the histone tails and for monitoring changes in the modifications in response to different stimuli. Antibodies against specific histone modifications (such as trimethylation of histone H3 lysine K27) are used for immunoprecipitation of that chromatin regions that have these modifications. ChIP can also be applied to study the binding of transcription factors and enzymes to the chromatin, using a specific antibody against the factor of interest. The associated chromatin regions can be analyzed either by PCR to detect specific loci or deep sequencing for a more global view, as well as by using microarray (ChIP-chip), but in the last case the resolution of data is weaker compared to sequencing. During the last decade the basic ChIP method has been developed which leads to the emergence of other diverse ChIP methods such as μChIP-seq for low, micro-scale sample amount analysis, and the modern single-molecule real-time sequencing (SMRT; “third generation sequencing”) of ChIP-samples.
Epigenetic modifications can also be studied by methods identifying the active regions of chromatin. In the DNase-seq method, the enzyme DNase I is utilized to digest the DNA that is not protected by nucleosome structure, so the regions that are sensitive to DNase are associated with active genes. Sequencing, microarray or Southern Blot techniques can be used for the results analysis and interpretation. The open chromatin can be detected also by the formaldehyde-assisted isolation of regulatory elements (FAIRE), which result in the isolation of open, nucleosome-depleted chromatin regions. The FAIRE method is based on the cross-linking between formaldehyde and DNA, histones and other proteins associated with it. The DNA is sonicated to be fragmented and after that isolated with phenol-chloroform extraction. Only DNA that is not bound by nucleosomes and associated proteins remains in the aqueous phase in the extraction, thus resulting in the isolation of the open and active regions of the genome. The isolated fragments can be analyzed again with different methods, such as PCR, microarray, and sequencing. Chromatin can also be studied by chromatin conformation capture (3C) technique that identifies the chromatin regions that are physically associated together, such as promoters with enhancers. 3C is often analyzed by PCR, but nowadays also by deep sequencing of the interactions globally (Hi-C).
Noncoding RNAs (ncRNAs) act as epigenetic modifiers to strongly regulate gene expression. Their aberrant expression mainly in the form of microRNAs (miRNAs) and long noncoding RNAs, may modify gene expression and trigger complicated immune disorders. The analysis of the RNA component of epigenetics most often is done on a genome-wide scale using next-generation sequencing methods. The changes in ncRNA and mRNA can be characterized by deep sequencing. The expression of RNAs can be analyzed also by quantitative Polymerase Chain Reaction (qPCR). Quantitative PCR also known as real-time PCR is a laboratory technique in molecular biology, by which the amount of the PCR product can be determined in real-time, and is very useful for investigating gene expression.
Transcriptomics
When genes are expressed, the genetic information stored in DNA is transferred to RNA (ribonucleic acid). RNAs are important macromolecules, composed of linear chains of nucleotides (Fig. 4), which are produced by the cellular process of transcription. RNAs perform diverse cellular and biological functions as either serve as templates for protein synthesis, or play critical catalytic and regulatory roles. In transcription the genetic information is transferred from DNA to mRNA. This process is carried out by an enzyme RNA polymerase. Several classes of RNAs exist in cells (messenger RNAs (mRNA), transfer RNAs (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), short interfering RNA (siRNAs), micro RNA (miRNAs), long non-coding RNA and pseudogenes), but those which take part in protein synthesis are messenger RNA (mRNA), transfer RNA (tRNA) and ribosomal RNA. Messenger RNA (mRNA) is a single-stranded molecule that mediate the transfer of genetic information to the ribosomes where the proteins are synthesized. In eukaryotes, each gene is transcribed to yield a single mRNA, whereas in prokaryotes, a single mRNA molecule may carry the genetic information from several genes; that is, several protein coding regions. A linear correspondence exists between the base sequence of a gene and the amino acid sequence of a polypeptide. Each group of three consecutive nucleotides encode the location of a particular amino acid in a protein molecule and each such triplet of bases is called a codon. Codons are translated into sequences of amino acids by ribosomes (which themselves consist of proteins and rRNA), tRNA, and helper proteins called translation factors.

The term “transcriptome” is widely used to designate the complete set of all the ribonucleic acid (RNA) molecules in a cell, tissue, or organism. The transcriptome reflects the molecular activity in cells, the genes that are actively expressed at any given moment. It encompasses all forms of RNAs molecules including protein coding, non-protein coding, alternatively spliced, alternatively polyadenylated, alternatively initiated, sense, antisense, and RNA-edited transcripts. Respectively, transcriptomics studies all types of transcripts within a cell or an organism, including mRNAs, miRNAs and different types of long noncoding RNAs (lncRNAs). Transcriptomics covers everything relating to RNAs such as their transcription and expression levels, functions, locations, trafficking, and degradation. It also includes the structures of transcripts and their parent genes with regard to start sites, 5′ and 3′ end sequences, splicing patterns, and posttranscriptional modifications.
The major studies of transcriptomics are aimed at:
- characterization of different states of cells (i.e. development stages), tissues or cell cycle phases by expression patterns;
- investigation of the molecular mechanisms underlying a phenotype;
- identifications of biomarkers differently expressed between the diseased state and healthy state;
- differentiation of disease stages or subtypes (e.g. cancer stages);
- setting up the causative relationship between genetic variants and gene expression patterns to illuminate the etiology of diseases.
During the last three decades the technological advance has revolutionized transcriptome profiling and redefined what is possible to investigate. Integration of transcriptomic data with other omics is giving an increasingly integrated view of cellular complexities facilitating holistic approaches to biomedical research.
The major techniques applied for transcriptome study are:
- Expressed sequence tag (EST)-based methods
- SAGE
- Hybridization-based microarray
- Real-time PCR
- NGS-based RNA-sequencing (RNA-seq) methods,
- RNA interference
- Bioinformatics tools for transcriptomes analysis.
The selection of the technique is dependent on cost effectiveness, sensitivity, high throughput, and minimal concentration of starting RNA. The methodology involves RNA isolation, purification, quantification, cDNA library construction, and high-throughput sequencing.
- Expressed sequence tag (EST)-based methods – Expressed sequence tags (ESTs) are relatively short DNA sequences (usually 200–300 nucleotides) usually generated from the 3′ ends of cDNA clones from which PCR primers can be derived and used to detect the presence of the specific coding sequence in genomic DNA. The sequencing of the ESTs gives an overview on the expression level of the gene. As more and more EST data have become publicly available, the usage of ESTs has expanded to other areas, such as in silico genetic marker discovery, in silico gene discovery, construction of gene models, alternative splicing prediction, genome annotation, expression profiling, and comparative genomics. In comparison with whole genome sequencing, EST technology is simpler and less costly, especially in the case of large genomes. There are EST databases such as dbEST (NCBI EST) (https://www.ncbi.nlm.nih.gov/genbank/dbest/) that contain sequence data and other information on “single-pass” cDNA sequences, or “Expressed Sequence Tags”, from a number of organisms and thus serve as a reference for the expression profile of an organism.
- Serial Analysis of Gene Expression (SAGE) is a transcriptomic technique used by scientists to produce a snapshot of the messenger RNA population in a sample of interest in the form of small tags that correspond to fragments of those transcripts. This technique is advantageous over EST because only short “tags” of about only 15 bases are sequenced. The short fragments generated are then joined together and sequenced. A pool of cDNA can be subjected to high-throughput NGS known as RNA-seq for quantification, discovery of novel ESTs, and profiling of RNAs.
- DNA chip, gene chip, biochip, or microarray is a collection of DNA, cDNA, oligonucleotides spots attached to a solid support such as glass or silicon chip. Through this hybridization-based method the expression levels of thousands of genes are monitored simultaneously. The major restriction of this technology is that genome sequence information is a prerequisite and also higher background inherent of hybridization technique.
- Quantitative real-time PCR (qRT-PCR) is a type of PCR for reliable quantification of low-abundance mRNA or low-copy transcripts. The major advantages of this technique are its high sensitivity, better reproducibility, and wide dynamic quantification range. It facilitates gene expressions and regulation studies even in a single cell based on its exponential amplification ability. The availability of diverse types of fluorescence monitoring system attached with the PCR resulted in its popularity for gene-expression studies. The problems of nonspecific amplification, formation of primer-dimers are some of the limitations of qRT-PCR.
- RNA-sequencing is one of the advanced high-throughput technology for transcriptomics. RNA-seq, also known as whole-transcriptome shotgun sequencing, utilizes NGS tools. The advantages of RNA-seq are that it does not rely on the availability of the genome sequence, has no upper quantification limits, shows high reproducibility, and possesses a large dynamic detection range.
The progress of transcriptomics led to a number of biological discoveries. This omic technology laid the foundation of the first real multi-omic studies, based on the comparisons between DNA sequence and mRNA expression. Nowadays, transcriptional analysis remains more frequently employed by most biologists since the data obtained are still more easily analyzed and shared than the more ‘downstream omics’ such as proteomics and metabolomics. More recently, transcriptomics is enjoying a second revival, as it is in many cases applicable to single cells.
Proteomics
Proteins command cellular structure and activity, provide the mechanisms for signaling between cells and tissues, and catalyze chemical reactions that support metabolism. The protein structure dictates the function (or dysfunction). Proteins can be the base cause of diseases (such as Alzheimer’s or Huntington’s disease), and they can be used to cure it (e.g., antibodies are used as therapeutics against viral and bacterial infections). The set of all expressed proteins in a biological system under a specific, defined conditions is known as proteome. The word “proteome” is a combination of protein and genome and was coined by Mark Wilkins in 1994.
The main aim of proteomics is a large-scale experimental analysis of the structure and function of this entire set of proteins produced by a living organism. The term “proteomics” first appeared in 1997 and referring to a core technology in systems biology approaches which studies the proteome. The function of cells is dependent on the proteins that are present in the intra- and intercellular space and their abundance. The synthesis of proteins in the cell is based on mRNA precursors but it is impossible to predict the abundance of specific proteins only based on analysis of gene expression. The reason for this is that native proteins undergo post-translational modifications (PTMs) or alterations in response to changes in the environment. The proteome is a dynamic reflection of both genes and the environment and is a valuable source for biomarker`s discovery because proteins are most likely to be ubiquitously affected in disease and disease response. The complete characterization of all proteins has been the goal of proteomics since its initiation almost 25 years ago. It tends to do more than merely identify proteins potentially present in a sample, but also to assess protein abundance, localization, posttranslational modifications, isoforms, and molecular interactions. Proteomics aims to study the flow of biological information through protein pathways and networks, with the eventual aim of understanding the functional relevance of proteins. This requires the development of technologies that can detect a wide range of proteins in samples from different origins. Various technologies are used in proteomics and most often they are applied in combination, for example one- or two-dimensional gel electrophoresis with mass spectrometry (MS) or liquid chromatography and MS.
The conventional techniques for isolation and purification of proteins are chromatography based such as ion exchange chromatography (IEC), hydroxyapatite, size exclusion chromatography (SEC) and affinity chromatography. For analysis of selective proteins, enzyme-linked immunosorbent assay (ELISA) and western blotting can be used. These techniques may be applied for analysis of individual proteins but they cannot define protein expression level. Sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE), two-dimensional gel electrophoresis (2-DE) and two-dimensional differential gel electrophoresis (2D-DIGE) techniques are used for separation of complex protein samples.
- Mass spectrometry (MS) enables the analysis of proteomes and usually is the preferred method for identifying proteins present in biological systems. The three primary applications of MS to proteomics are: cataloging protein expression, defining protein interactions, and identifying sites of protein modification. Mass spectrometry measures the mass-to-charge ratio (m/z) of gas-phase ions. Mass spectrometers consist of an ion source that converts analyte molecules into gas-phase ions, a mass analyzer that separates ionized analytes based on m/z ratio, and a detector that records the number of ions at each m/z value. The development of electrospray ionization (ESI) and matrix-assisted laser desorption/ionization (MALDI), the two soft ionization techniques capable of ionizing peptides or proteins, revolutionized protein analysis using MS. The mass analyzer is central to MS technology. For proteomics research, four types of mass analyzers are commonly used: quadrupole (Q), ion trap (quadrupole ion trap, QIT; linear ion trap, LIT or LTQ), time-of-flight (TOF) mass analyzer, and Fourier-transform ion cyclotron resonance (FTICR) mass analyzer. They vary in their physical principles and analytical performance. “Hybrid” instruments have been designed to combine the capabilities of different mass analyzers and include the Q-Q-Q, Q-Q-LIT, Q-TOF, TOF-TOF, and LTQ-FTICR.
- Tandem mass spectrometry is additional experimental procedure which is also known as MS/MS or MS2. This is a key technique for protein or peptide sequencing and PTM analysis where two or more mass analyzer are coupled. The molecules of a given sample are ionized and the first spectrometer (designated MS1) separates these ions by their mass-to-charge ratio (often given as m/z or m/Q). Ions of a particular m/z-ratio coming from MS1 are selected and then made to split into smaller fragment ions, e.g. by collision-induced dissociation (CID), electron-capture dissociation (ECD), electron-transfer dissociation (ETD), ion-molecule reaction, or photo dissociation. These fragments are then introduced into the second mass spectrometer (MS2), which in turn separates the fragments by their m/z-ratio and detects them. The fragmentation step makes it possible to identify and separate ions that have very similar m/z-ratios in regular mass spectrometers. The complexity of the biological systems requires that the proteome be separated before analysis. Both gel chromatography- and liquid chromatography-based separations have proven useful in this regard. Typically, after these extensive separations, proteins are characterized by MS analysis of either intact protein (top–down) or enzymatically digested protein peptides (bottom–up). Protein identifications are made by comparing measured masses of intact proteins (top–down) or digested protein peptides (bottom–up) to calculated masses obtained from genome data.
- Isobaric Taq Labelling (ITL) is a commonly used method for quantification of proteins. In proteomics the quantification of the protein abundance is an important focus. Protein expression levels represent the balance between translation and degradation of proteins in cells. It is therefore assumed that the abundance of a specific protein is related to its role in cell function. ITL provides opportunity for simultaneous identification and quantification of proteins from multiple samples in a single analysis. To measure proteins` quantity in a given sample, peptides are labeled with chemical tags that have the same structure and nominal mass, but vary in the distribution of heavy isotopes in their structure. These tags, commonly referred to as tandem mass tags, are designed so that the mass tag is cleaved at a specific linker region upon higher-energy collisional-induced dissociation (HCD) during tandem mass spectrometry yielding reporter ions of different masses. Protein quantitation is accomplished by comparing the intensities of the reporter ions in the MS/MS spectra. Other quantitative techniques are ICAT labelling and Stable Isotopic Labeling with Amino Acids in Cell Culture (SILAC). The ICAT has also expanded the range of proteins that can be analyzed and permits the accurate quantification and sequence identification of proteins from complex mixtures. SILAC is an MS-based approach for quantitative proteomics that depends on metabolic labeling of whole cellular proteome. The proteomes of different cells grown in cell culture are labeled with “light” or “heavy” form of amino acids and differentiated through MS.
- X-ray crystallography and nuclear magnetic resonance (NMR) spectroscopy are two major high-throughput techniques that provide three-dimensional (3D) structure of protein that might be helpful to understand its biological function. With the support of high-throughput technologies, a huge volume of proteomics data is collected. Bioinformatics databases are established to handle enormous quantity of data and its storage. Various bioinformatics tools are developed for 3D structure prediction, protein domain and motif analysis, rapid analysis of protein–protein interaction and data analysis of MS. The alignment tools are helpful for sequence and structure alignment to discover the evolutionary relationship. Proteome analysis provides the complete depiction of structural and functional information of cell as well as the response mechanism of cell against various types of stress and drugs using single or multiple proteomics techniques.
The major techniques used in the proteomics research are given in the Fig. 5.

Metabolomics
The metabolome was first described by Oliver and colleagues in 1998, during their pioneering work on yeast metabolism, as the complete complement of small molecules in a biological system or fluid. It includes all the small molecules such as lipids, amino acids, fatty acids, carbohydrates and vitamins which are known as metabolites and are produced as a result of cellular metabolism. They participate in the metabolic processes in the cells by interacting with other biological molecules following metabolic pathways. The metabolites` status in the biological systems is highly variable and time-dependent and changes in levels of key metabolites are observed due to the genetic, environmental, nutritional and other factors.
Metabolomics (metabolome analysis) is the most recently introduced strategy among the Omics that systematically identify and quantify the metabolites present in a cell, tissue, organ, biofluids, or organism at a specific point of time. Methods used in metabolomics aim to measure low molecular weight compounds (metabolites) with different physical characteristics such as compound polarity, functional groups and structural similarity. On the basis of these properties the metabolome is divided into subsets of various metabolites and for their investigation analytical procedures optimized for each type of molecule is applied in metabolomics.
The metabolites are products of biochemical reactions which are carried out with the participation of proteins of the proteome (Fig. 6). Their concentrations are dependent on the interactions of other processes (transcription, translation, cellular signaling, etc.), they have been investigated as reporters for metabolism in PD. This in turn determines the biological structure and function of the final phenotype of the organism. Therefore, changes in gene expression, proteins` function and the environment directly affect the metabolites` concentration in a biological system. The metabolome is composed of a relatively small number (around 5000), but different types of metabolites which make the metabolome more physically and chemically complex than genome, transcriptome and proteome. Moreover, the genome, transcriptome and proteome consist of compounds which are part of the metabolome. In addition, the components of the metabolome are highly conserved between organisms in comparison to genome, transcriptome and proteome which is why it is considered that the metabolome is evolutionary the oldest part of the cell.

Metabolomics is the most appropriate strategy for determination of disease-associated biomarkers. Comprehensive study of metabolites is a desirable tool for diagnosing disease, identifying new therapeutic targets, and enabling appropriate treatments. The analyses of small numbers of metabolites are used in diseases` diagnostics for decades, for example, the development of blood glucose test strips in the 1950s to test for diabetes or quantifying phenylalanine in newborns to screen for phenylketonuria.
Metabolomics analyses can be grouped into two main approaches: targeted and untargeted metabolomics. In targeted analysis the metabolites which are determined are known and represent specific pathway(s) or class(es) of molecules. On the contrary, the untargeted metabolomics aims to quantify and identify as many metabolites as possible. Targeted metabolomics refers to absolute quantification (nM or mg/ml) and uses an internal standard and semi quantitative or quantitative analysis to detect known compounds related to specific pathways. The untargeted approach measures all the present metabolites in a sample and apply relative quantification (fold change) and comparison between samples. The key role that metabolomics plays in multi-omics integration and system modelling is due to the fact that it can be quantitative. Systems modeling cannot be performed without accurate values or accurate concentrations as inputs and, likewise, systems models cannot be easily verified without accurate, quantitative concentrations as outputs. Metabolomics can deliver both (quantitative input and output data), making it extremely valuable to systems modelers.
Due to the metabolome`s enormous chemical complexity, it cannot be studied comprehensively by a single technology. The earliest application of metabolomics dates back to 1970s when gas chromatography–mass spectrometry was used for metabolite profiling of clinical urine samples. Nicholson and others in the 1980s made a profile of clinical samples by the application of nuclear magnetic resonance spectroscopy. Mass spectrometry coupled with both liquid chromatography and capillary electrophoresis are also applied in metabolomics to complement the analytical techniques available. Some of the analytical instruments utilized for the purposes of metabolomics research are being applied frequently and routinely while others have specific roles and are applied less frequently.
- Gas chromatography – Mass spectrometry (GC-MS) – Among the commonly used methods in metabolomics is GC-MS. As the name implies, GC-MS unifies two techniques to form a single method for analyzing mixtures of chemicals. Gas chromatography separates the components of a mixture and mass spectroscopy characterizes each of the components individually. By combining the two techniques, each metabolite in a sample can be both qualitatively and quantitatively evaluate. This chromatographic technique is used to study metabolites which have a low boiling point and which will be present in the gas phase at the temperature range 50-350°C. These metabolites can have a low boiling point in their biologically native form or the boiling point of a metabolite can be decreased through a chemical alteration, also known as chemical derivatisation. The sample containing metabolites is introduced (injected) into a mobile phase, which in the gas chromatography is an inert gas such as helium. The mobile phase carries the sample mixture through what is referred to as a stationary phase. The metabolites are separated on the basis of their adsorption to the stationary phase. The stationary phase is usually contained in a glass or stainless steel column and represents a chemical that can selectively adsorb components in a sample mixture. By changing characteristics of the mobile phase and the stationary phase, different mixtures of chemicals can be separated. As the individual metabolites elute from the GC column, they enter the electron ionization (mass spec) detector. There, they are bombarded with a stream of high-energy electrons (70 eV) which causing them to break apart into fragments. These fragments can be large or small pieces of the original molecules. The mass spectrum obtained for a given chemical compound is basically the same every time. Therefore, the mass spectrum is a fingerprint for the molecule. This fingerprint can be used to dentify the compound.
- Liquid Chromatography-Mass Spectrometry (LC-MS) – The main difference between GC-MS and LC-MS is that in liquid chromatography (LC), the mobile phase is a solvent. The LC separates mixture of metabolites which are in liquid form, usually contains methanol, acetonitrile and water. By using different packing of columns (different stationary phase) with high efficiency small amount of complex mixture can be separated. The columns used in HPLC are consists of Octadecyl (C18), Octyl (C8), Cyano, Amino, Phenyl packing’s and generally their length is about 50mm to 300mm. The columns are used on the basis of nature of compounds to be separated. This liquid containing mixture of components is transferred into the ion source of mass spectrometer where the process of ionization takes place and droplets carry an excess of positive or negative electric charge are formed. The ionization could be achieved by the use of different types of ionization sources and interfaces. After ionization the ions are transferred into mass Analyser where the separation of ions are done according to their mass to charge (m/z) ratio.
- Capillary Electrophoresis-Mass Spectrometry (CE-MS) – this analytical technique is also called capillary zone electrophoresis-mass spectrometry. CE-MS involves the separation of ionic species in the liquid phase via the application of high voltages and is usually coupled to an electrospray mass spectrometry. Capillary electrophoresis separates metabolites based on their electrophoretic mobility in a liquid electrolyte solution operating in an electric field. Electrophoretic mobility is dependent on the charge and size of the metabolite and so separation of metabolites with different sizes and/or charges is possible. All metabolites have to be charged to allow any mobility to occur. In GC-MS, LC-MS and CE-MS there is a separation of metabolites before their detection with mass spectrometry. This peculiarity provides the ability to detect metabolites at low concentrations, commonly nanomoles/litre (nM/L) or micromoles/litre (µM/L).
- Nuclear magnetic resonance (NMR) spectroscopy is the most widely used analytical instruments in metabolomics research together with Mass spectrometry. NMR spectroscopy applies the magnetic properties of atomic nuclei in a metabolite. Only some atoms are NMR active and include 1H, 13C and 31P; proton (1H). The technique operates by placing a liquid sample in a small internal diameter tube (for example a 5 mm tube), or occasionally a piece of tissue is studied directly using a special sample holder. The sample is pulsed with a range of radio frequencies covering all possible energies required for exciting the selected type of nuclei. The nuclei absorb energy at different radio frequencies depending on their chemical environment and then the release of this energy is measured, forming a free induction decay (FID). This FID is converted from a time domain data set to a frequency domain – using a Fourier transformation – and an NMR spectrum is constructed as the chemical shift (effectively the absorption energy) plotted against peak intensity.
- Vibrational spectroscopy techniques such as Fourier transform–infrared (FT-IR) spectroscopy and Raman spectroscopy are applied for analysis of the metabolic changes in biological samples. The techniques` principle is based on the ultraviolet or infrared light transmission through a sample, or sometimes reflection of the light from the sample, before its detection. The approaches predominantly measure the vibrations and rotations of bonds related to different chemical functional groups resulting from the interaction of the sample with the ultraviolet or infrared light. These techniques usually lack the specificity to detect each metabolite separately but instead specific parts of the molecule absorb the ultraviolet or infrared light at specific wavelengths. For example, in FT-IR, C-H stretching vibrations characteristic of fatty acid chains are observed between the wavenumber range 3100 – 2800 cm-1 and the region between 1800 and 1500 cm-1 is dominated by amide I and amide II bands indicating the predominance of either alpha helix or beta sheet structures.
In the field of metabolomics, many hundreds of samples are routinely analyzed, and a minimum of several hundreds of metabolites are usually detected. The data obtained from the detection of metabolites in a given biological samples is followed by further analysis using statistical multivariate methods (chemometrics) to extract biological, physiological and clinically relevant information.
Advantages and disadvantages of omics technologies
The implementation of omics technologies and the integration of omics data has been realized for a broad range of research areas, including food and nutrition science, systems microbiology, analysis of microbiomes, genotype–phenotype interactions, systems biology, natural product discovery and disease biology. Recently omics-based approaches have been significantly improved with the addition of novel concepts such as exposome/exposomics (role of the environment in human diseases), adductomics (study of compounds that bind DNA and cause damage and mutations), volatilomics (study of volatile organic compounds to the metabolomics/lipidomics analysis), nutrigenomics (study of how foods affect our genes), etc. However, omics technologies are aimed at primarily four omics research fields – genomics, transcriptomics, proteomics and metabolomics. The traditional research fields like genetics, pharmacogenetics and toxicogenetics which do not apply the omics concept, provide the static sequences of genes and proteins. On the contrary, the omics technologies provide simultaneously measurement of the number of proteins, genes, metabolites and enables to get large scale data in a short time. The collective analysis of the biological processes in a biological system provided by all omics is an advantage to previous traditional methods in better understanding of the whole picture of the biological system. Although large-scale omics data are becoming more accessible, the integration of real multi-omics data is still very challenging. This is due to fact that many of the specific analycal tools and experimental designs that are conventionally used for individual omics disciplines (genomics, transcriptomics, and proteomics) are not sufficiently well-suited to allow reliable comparisons and integration across multiple omics disciplines. For example, methods for sample collection and storage and the quantity and type of biological samples required in genomics studies are often not compatible with metabolomics, proteomics or transcriptomics. During collecting data, there are many problems such as data heterogeneity, small sample size in comparison to many parameters, confirmation and interpretation of data due to many interactions in biological system and deficient information on those systems. Almost in all omics experiments hundreds to thousands of target molecules and variables are measured (metabolites, proteins, genes, transcripts and SNPs). This requires the investigation of sufficient number of biological samples in aim to provide statistically reliable results. Moreover, the multi-omics data must be generated from the same set of samples to permit the direct comparison under the same conditions. However, this is not always possible due to limitations in sample biomass, sample access or financial resources. In single omics experiments, larger sample sizes are often required to overcome this problem. In addition, the integrated multi-omics data must be analyzed as single data sets before being deposited into omics-specific databases in order to make it publicly available. These issues emphasize that for the high-quality large scale-omics studies is necessary to : 1) plan an experiment properly, 2) collect, prepare, and store biological samples attentively, 3) collect quantitative multi-omics data and associated meta-data thoughtfully, 4) utilize proper tools for integration and interpretation of the data.
Some of the advantages and disadvantages of the main omics technologies are summarized in Table 1.
| Omics technology | Advantage | Disadvantage |
|---|---|---|
| Genomics | • Analysis of the complete genome • Studying gene polymorphism among individuals • Through SNPs identification a valuable information for early diagnosis and treatment of disease is provided • Easy to implement lab techniques |
• Performed genome analysis is not enough to predict the final biological effect of DNA due to epigenetics, post-transcriptional and posttranslational changes. • Require specific and expensive lab equipment |
| Epigenomics | Provide knowledge about regulation of gene expression | • It is difficult to relate the obtained epigenetic data with gene expression, because the transcription can be affected by multiple processes • Not every methylated region can be detected with the techniques applied. |
| Transcriptomics | • Study of the complete set of mRNA (transcriptome) • Identification of the major pathways involved in drug response and toxicity. • Easy to implement lab techniques |
• Protein expression is influenced by post-translational changes which lead to incorrect data • Require expensive and specific equipment |
| Proteomics | • Study of the complete set of proteins in a biological system • Detection of unknown and unexpected proteins |
• Expensive equipment and time-consuming procedures not applicable to entire proteome • Some proteins are difficultly separated and purified • MS application and interpretation of data require specially trained staff |
| Metabolomics | • Provides adequate data about modifications in the metabolic processes in the cell • Endogenous metabolites are less than genes, transcripts and proteins, so fewer data have to be interpreted • Detecting of diseases` biomarkers |
• Relatively low numbers of metabolites (a few thousand), as can be measured • Specific, expensive equipment • The use of MS and NMR require specially trained staff |
Multi-omics datasets can provide a greater depth of understanding in certain scenarios, but this is not without cost. These studies frequently are based on a large numbers of comparisons, correct data type, relevant statistical analyses, specific equipment and skilled personnel and considerable investment of time and money. When designing an experiment it must be taken into account what types of omics data can and should be integrated to obtain powerful biological insights of the system being studied. The application of high throughput omics platforms is not always necessary to answer the research question. The results obtained from the traditional techniques, such as quantitative polymerase chain reaction (qPCR), enzyme-linked immunosorbent assay (ELISA), immunohistochemistry (IHC), sometimes may be enough to provide understanding for biological mechanisms. These techniques often complement the findings from a larger omics study as they are applied to verify the significant molecule identified from omics data is a true positive result.
Test: LO1- Basic level
References
- Altaf-Ul-Amin M, Afendi FM, Kiboi SK, Kanaya S. 2014. Systems Biology in the Context of Big Data and Networks, BioMed Res Int, 1:11.
- Aslam B, Basit M, Nisar MA, Khurshid M, Rasool MH. 2017. Proteomics: Technologies and Their Applications. J Chromatogr Sci, 55(2).
- Benavente L, Goldberg A, Myburg H, Chiera J, Zapata M & van Zyl L. 2015. The application of systems biology to biomanufacturing. Pharm. Bioprocess, 3(4):341–355.
- Bertolaso M, Giuliani A, de Gara L. 2010. Systems Biology Reveals Biology of systems. Wiley Periodicals, Inc., 16 (6).
- Bloom FE. 2001. What does it all mean to you? J Neurosci, 21:8304–8305.
- Clish CB. 2015. Metabolomics: an emerging but powerful tool for precision medicine. Cold Spring Harb Mol Case Stud, 1: a000588
- Coorssen JR. 2013. Proteomics. Brenner’s Encyclopedia of Genetics (Second Edition), p. 508-510.
- Dos Santos EAF, Santa Cruz EC, Ribeiro HC, Barbosa LD, Zandonadi FS, Sussulini A. 2020. Multi-omics: An Opportunity to Dive into Systems Biology. Braz J Anal Chem, 7(29): 18-44.
- Haas R, Zelezniak A, Iacovacci J, Kamrad S, Townsend S, Ralser M. 2017. Designing and interpreting ‘multi-omic’ experiments that may change our understanding of biology. Curr Opin Syst Biol, 6:37–45.
- Han X, Aslanian A, Yates JR. 2008. Mass Spectrometry for Proteomics. Curr Opin Chem Biol, 12(5): 483–490.
- Horgan RP, Kenny LC. 2011. SAC review ‘Omic’technologies: genomics, transcriptomics, proteomics and metabolomics. Obstet Gynecol, 13:189–195
- Ideker T, Thorsson V, Ranish JA, ChristmasR, Buhler J, Eng R, Bumgarner JK, Goodlett DR, Aebersold R, Hood L. 2001. Integrated genomic and proteomic analyses of a systematically perturbed metabolic network. Science, 292:929–934.
- Ishii N, Tomita M. 2009. Multi-Omics Data-Driven Systems Biology of E. coli. In: Lee S.Y. (eds) Systems Biology and Biotechnology of Escherichia coli. Springer, Dordrecht.
- Kalavacharla V, (Kal), Subramani M, Ayyappan V, et al. 2017. Plant Epigenomics. Handbook of Epigenetics, 245–258.
- Karahalil B. 2016. Overview of Systems Biology and Omics Technologies. Curr Med Chem, 23:1-10.
- Karamperis K, Wadge S, Koromina M. 2020. Genetic testing. Applied Genomics and Public Health (Translational and Applied Genomics), Elsevier, 189-207.
- Laurence P. 2015. The case of the gene: Postgenomics between modernity and postmodernity. EMBO Reports, 16 (7): 777–781.
- Liang KH. 2013. Transcriptomics. Bioinformatics for Biomedical Science and Clinical Applications. Woodhead Publishing, p. 49-82.
- Lin, S., Fang, L., Li, C., Liu, G. 2019. Epigenetics and heritable phenotypic variations in livestock. In: Tollefsbol T. (editor). Transgenerational Epigenetics. 2nd edition. Amsterdam, Netherlands: Elsevier, p. 283-313
- Milward EA, Shahandeh A, Heidari M, Johnstone DM, Daneshi N, Hondermarck H. 2016. Transcriptomics. Encyclopedia of Cell Biology. Elsevier, The Netherland, 4: 160-165.
- O’Donnell ST, Ross RP, Stanton C. 2020. The Progress of Multi-Omics Technologies: Determining Function in Lactic Acid Bacteria Using a Systems Level Approach. Front. Microbiol, 10: 1-17.
- Pereira C, Adamec BJ. 2019. Metabolome Analysis. Encyclopedia of Bioinformatics and Computational Biology, 3:463-475
- Pinu FR, Beale DJ, Paten AM, Kouremenos K, et al. 2019. Systems Biology and Multi-Omics Integration: Viewpoints from the Metabolomics Research Community. Metabolites, 9(76).
- Potters, G. 2010. Systems Biology of the Cell. Nature Education, 3(9):33.
- Pratima NA, Gadikar R. 2018. Liquid Chromatography-Mass Spectrometry and Its Applications: A Brief Review. Arc Org Inorg Chem Sci, 1(1): 26-34.
- Shah TR, Ambikanandan M. 2011. Proteomics. Challenges in Delivery of Therapeutic Genomics and Proteomics. Elsevier, p. 387-427.
- Strange K. 2005. The end of “naïve reductionism”: rise of systems biology or renaissance of physiology? Am J Physiol Cell Physiol, 288: C968–C974.
- Turunen TA, … Ylä-Herttuala S. 2018. Epigenomics. Encyclopedia of Cardiovascular Research and Medicine, p. 258-265.
- Vlaanderen J, Moore LE, Smith MT, et al. 2010. Application of omics technologies in occupational and environmental health research: current status and projections. Occup Environ Med, 67: 136-143.
- Yadav D, Tanveer A, Malviya N, Yadav S. 2018. Overview and Principles of Bioengineering: The Drivers of Omics Technologies and Bio-Engineering Towards Improving Quality of Life, p.3-23
- Yang X. 2020. Multitissue Multiomics Systems Biology to Dissect Complex Diseases. Trends Mol Med, 26 (8).
- A brief guide to genomics. https://www.genome.gov/about-genomics/fact-sheets/A-Brief-Guide-to-Genomics
- Analytical Techniques applied in Metabolomics. https://www.futurelearn.com/info/courses/metabolomics/0/steps/10710
- Boundless biology. https://courses.lumenlearning.com/boundless-biology/chapter/the-genetic-code/
- Altaf-Ul-Amin M, Afendi FM, Kiboi SK, Kanaya S. 2014. Systems Biology in the Context of Big Data and Networks, BioMed Res Int, 1:11.
- Aslam B, Basit M, Nisar MA, Khurshid M, Rasool MH. 2017. Proteomics: Technologies and Their Applications. J Chromatogr Sci, 55(2).
- Benavente L, Goldberg A, Myburg H, Chiera J, Zapata M & van Zyl L. 2015. The application of systems biology to biomanufacturing. Pharm. Bioprocess, 3(4):341–355.
- Bertolaso M, Giuliani A, de Gara L. 2010. Systems Biology Reveals Biology of systems. Wiley Periodicals, Inc., 16 (6).
- Bloom FE. 2001. What does it all mean to you? J Neurosci, 21:8304–8305.
- Clish CB. 2015. Metabolomics: an emerging but powerful tool for precision medicine. Cold Spring Harb Mol Case Stud, 1: a000588
- Coorssen JR. 2013. Proteomics. Brenner’s Encyclopedia of Genetics (Second Edition), p. 508-510.
- Dos Santos EAF, Santa Cruz EC, Ribeiro HC, Barbosa LD, Zandonadi FS, Sussulini A. 2020. Multi-omics: An Opportunity to Dive into Systems Biology. Braz J Anal Chem, 7(29): 18-44.
- Haas R, Zelezniak A, Iacovacci J, Kamrad S, Townsend S, Ralser M. 2017. Designing and interpreting ‘multi-omic’ experiments that may change our understanding of biology. Curr Opin Syst Biol, 6:37–45.
- Han X, Aslanian A, Yates JR. 2008. Mass Spectrometry for Proteomics. Curr Opin Chem Biol, 12(5): 483–490.
- Horgan RP, Kenny LC. 2011. SAC review ‘Omic’technologies: genomics, transcriptomics, proteomics and metabolomics. Obstet Gynecol, 13:189–195
- Ideker T, Thorsson V, Ranish JA, ChristmasR, Buhler J, Eng R, Bumgarner JK, Goodlett DR, Aebersold R, Hood L. 2001. Integrated genomic and proteomic analyses of a systematically perturbed metabolic network. Science, 292:929–934.
- Ishii N, Tomita M. 2009. Multi-Omics Data-Driven Systems Biology of E. coli. In: Lee S.Y. (eds) Systems Biology and Biotechnology of Escherichia coli. Springer, Dordrecht.
- Kalavacharla V, (Kal), Subramani M, Ayyappan V, et al. 2017. Plant Epigenomics. Handbook of Epigenetics, 245–258.
- Karahalil B. 2016. Overview of Systems Biology and Omics Technologies. Curr Med Chem, 23:1-10.
- Karamperis K, Wadge S, Koromina M. 2020. Genetic testing. Applied Genomics and Public Health (Translational and Applied Genomics), Elsevier, 189-207.
- Laurence P. 2015. The case of the gene: Postgenomics between modernity and postmodernity. EMBO Reports, 16 (7): 777–781.
- Liang KH. 2013. Transcriptomics. Bioinformatics for Biomedical Science and Clinical Applications. Woodhead Publishing, p. 49-82.
- Lin, S., Fang, L., Li, C., Liu, G. 2019. Epigenetics and heritable phenotypic variations in livestock. In: Tollefsbol T. (editor). Transgenerational Epigenetics. 2nd edition. Amsterdam, Netherlands: Elsevier, p. 283-313
- Milward EA, Shahandeh A, Heidari M, Johnstone DM, Daneshi N, Hondermarck H. 2016. Transcriptomics. Encyclopedia of Cell Biology. Elsevier, The Netherland, 4: 160-165.
- O’Donnell ST, Ross RP, Stanton C. 2020. The Progress of Multi-Omics Technologies: Determining Function in Lactic Acid Bacteria Using a Systems Level Approach. Front. Microbiol, 10: 1-17.
- Pereira C, Adamec BJ. 2019. Metabolome Analysis. Encyclopedia of Bioinformatics and Computational Biology, 3:463-475
- Pinu FR, Beale DJ, Paten AM, Kouremenos K, et al. 2019. Systems Biology and Multi-Omics Integration: Viewpoints from the Metabolomics Research Community. Metabolites, 9(76).
- Potters, G. 2010. Systems Biology of the Cell. Nature Education, 3(9):33.
- Pratima NA, Gadikar R. 2018. Liquid Chromatography-Mass Spectrometry and Its Applications: A Brief Review. Arc Org Inorg Chem Sci, 1(1): 26-34.
- Shah TR, Ambikanandan M. 2011. Proteomics. Challenges in Delivery of Therapeutic Genomics and Proteomics. Elsevier, p. 387-427.
- Strange K. 2005. The end of “naïve reductionism”: rise of systems biology or renaissance of physiology? Am J Physiol Cell Physiol, 288: C968–C974.
- Turunen TA, … Ylä-Herttuala S. 2018. Epigenomics. Encyclopedia of Cardiovascular Research and Medicine, p. 258-265.
- Vlaanderen J, Moore LE, Smith MT, et al. 2010. Application of omics technologies in occupational and environmental health research: current status and projections. Occup Environ Med, 67: 136-143.
- Yadav D, Tanveer A, Malviya N, Yadav S. 2018. Overview and Principles of Bioengineering: The Drivers of Omics Technologies and Bio-Engineering Towards Improving Quality of Life, p.3-23
- Yang X. 2020. Multitissue Multiomics Systems Biology to Dissect Complex Diseases. Trends Mol Med, 26 (8).
- A brief guide to genomics. https://www.genome.gov/about-genomics/fact-sheets/A-Brief-Guide-to-Genomics
- Analytical Techniques applied in Metabolomics. https://www.futurelearn.com/info/courses/metabolomics/0/steps/10710
- Boundless biology. https://courses.lumenlearning.com/boundless-biology/chapter/the-genetic-code/
Green Energy & ICT: From Smart to Wise Strategies
B A S I C L E V E L
Nowadays, new regulations are in preparation aiming to propose contact code information for specific fields and also for individuals to find alternative sources of energy.
The Green energy
Nowadays, new regulations are in preparation aiming to propose contact code information for specific fields and also for individuals to find alternative sources of energy. These alternative sources have to provide power for private and public buildings, and at the same time generate a low number of toxic compounds. The is the so-called” going green” approach. Different alternative types of energy have been designed, like solar and nuclear power, with the purpose to save the planet because the toxic emissions accompanying the production of the traditional ones are a huge problem since they affect badly world life.
Along with the widely investigated toxic actions of global warming in recent years, it is vague in parallel to the harm provided by other resources used in the production of food and maintenance of clean water. Society has referred to materials like coal, oil, and even kerosene to ensure the needed energy. The fossil fuels, coal, oil, and other resources used for power production emit harmful side effects. These fuels are non-renewable and contaminate the environment and the atmosphere, impacting the sources needed for the survival of the species inhabiting our planet. As these sources are naturally restricted, troubles for their shortages and access are growing, the worst among them is their harmful effect on the environment. So, the use of these conventional sources of energy contributes to global warming. Coal and oil outflow toxic gases into the environment, and in this way endanger general health rising respiratory problems, and diminish the quality of life.
Green energy will help to relieve and smoothen at least some of these problems, and the faster we move to renewable energy sources the better.
What is green energy?
Biosensors are sensors that transform bio-recognition processes through a physico-chemical transducer into observable signals, with electronic and optical techniques as two main transducers. The creation of biosensors addresses today’s rapidly rising need for clinical diagnostics. A combination of advantages is brought on by the use of biosensors. Biosensors, first, are highly sensitive. This is because biomolecules have a high affinity for their targets, for example, antibodies catch antigens with a dissociation constant at the nanomolar scale, and DNA – DNA interactions are so much stronger than antigen-antibody. Second, biological recognition is typically very selective. The enzyme and substrate are much like a lock and a key, for example. Such high selectivity frequently leads to biosensors that are selective. Third, the production of inexpensive, integrated, and ready-to-use biosensor devices has become relatively easy to develop due to the development of the modern electronic industry. The ability to detect pathogens or perform genetic analysis in hospitals is certainly improved by these biological sensors; more importantly, they are especially useful for small clinics and even point-of-care analysis.
For biosensors with clinical applications, a range of new techniques have been developed. Biosensors are, in general, analytical devices constructed of an element of biological recognition and an optical/electronic transducer. The biological element is responsible for the capture of solution analytes and the transducer transforms the binding event to a measurable signal variation. By the nature of recognition, enzyme-based biosensors, immunological biosensors, and DNA biosensors, could categorize the type of biosensors. In addition, electronic biosensors (electrical or electrochemical), optical biosensors (fluorescent, surface plasmon resonance, or Raman), and piezoelectric biosensors (quartz crystal microbalance) are available depending on the type of transducer.
Green energy products work
It is accepted that to be recognized as a green energy resource it should not produce pollution, such as is found with fossil fuels. This evidence indicates that not all sources used by the renewable energy industry are green. Thus, power generation that burns organic material from sustainable forests is mind as renewable but due to the CO2 production by the burning process, it is not green. Green energy sources are usually naturally completed, as set against fossil fuel sources like natural gas or coal, which are developed millions of years. Green sources also often avoid mining or drilling operations which can impair the ecosystems.
The future in energy consumption is focused on the exploitation of a mix of green, renewable, and conventional energy, regardless of the purchased product. Thus, all energy sources in the electric grid are mixed alongside the power transmission grid. For those keen to become green at home and do not possess opportunities for a solar panels array, the mentioned mix is the best way to reduce the carbon footprint associated with energy consumption. It is the most presumable way to rise the large-scale renewable energy investment and it gives more households and businesses access to green energy.
The types of green energy
The variety of green energy types are linked with the wide variety of sources (Fig. 1). Some of these types are better suited to specific environments or regions. As a source of energy, green energy often comes from renewable energy technologies such as solar energy, wind power, geothermal energy, biomass, and hydroelectric power. These technologies are acting through different processes.
Find an alternative energy source is the focus of many countries around the world. It is of strategic importance to discover natural and renewable options as an energy source. In some cases, it may be a simple decision to find appropriate architectural designs that keep buildings cool in the summer and warm in the winter. In another, an anaerobic design is used in energy-producing systems to replace fossil fuels with other resources. Thus, reducing carbon emissions, preventing environmental harm, and jobs-creating are just some of the advantages provided by investing in green energy.


Figure 1. The variety of green energy types
Going green means greater funding to solar, wind, and other renewable energy projects, creating technologies to better harness the renewable sources and make them more admissible for people.
Energy Consumption in ICT Sector
Currently, it is estimated that ICT consumes 1.15% of the total electricity supply. It is accounted that the total annual operational electricity consumption for ICT is 242 TWh in 2015. This sum comprises on-site generated electricity (27 TWh) and grid electricity (215 TWh). Besides, the global operational carbon emitted by the ICT sector in 2015 is about 169 M tonnes CO2. It is equivalent to 0.53% of the whole carbon emission by the energy sector (32 G tonnes) and 0.34% of global carbon emission (50 G tonnes) in 2015. The electricity consumption in the ICT network increased by 31% from 2010 to 2015. This subsummes185 TWh rise for 5 years, which is corresponding to 1% of the total electricity grid supply. The operational carbon emission growth has been 17% for this period.
Applying widely the 5G in near future, the rate of energy consumption increase is going to be even greater.
Green Energy Provisioning for ICT
Today, reducing greenhouse gas (GHG) emissions is getting up one of the most important research subjects in Information and Communication Technologies (ICT) due to the disturbing growth of indirect GHG emissions coming from the tremendous use of ICT electrical devices. Thus, solving the ICT GHG problem is linked to improving energy efficiency through energy consumption reduction at the micro-level.
Research in this field is focused on microprocessor design, computer design, power-on-demand architectures, and virtual machine consolidation techniques. The micro-level energy efficiency method will bring an overall rise in energy consumption due to the Khazzoom–Brookes postulate (also known as Jevons paradox) that says: “energy efficiency improvements that, on the broadest considerations, are economically justified at the micro-level, lead to higher levels of energy consumption at the macro-level”. Therefore, it could be reasonable that reducing GHG emissions at the macro level is a more appropriate solution. Large ICT companies, like Microsoft which consumes up to 27MW of energy at any given time, have built their data centers near green power sources. Unfortunately, a lot of computing centers are not placed close to green energy sources. For this reason, green energy distributed network is an emerging technology, allowing that losses incurred in energy transmission over power utility infrastructures are much higher than those caused by data transmission. This makes relocating a data center near a renewable energy source a more efficient decision unless bringing the energy to an existing location.
Management and technical policies will be a decision to arrange virtualization, which helps to move virtual infrastructure resources from one site to another based on power availability. This will facilitate the use of renewable energy within the ICT network providing an ‘Infrastructure as a Service (IaaS)’ management tool.
Green ICT and ways of Greening ICT
Green ICT is a broad concept, which lacks a general definition. This concept includes the energy efficiency of equipment such as computers, servers, and monitors. Occasionally, it mentioned the production of ICT equipment as well as the recycling. In other cases, it includes ways in which ICT can be used to mitigate the environmental impact of other sectors.
- Defining Green ICT
Green ICT has been defined in the literature as “the using of IT resources in an energy-efficient and cost-effective manner” or “an initiative to encourage individuals, groups, and organizations engaged in the use of ICT to consider environmental problems and find solutions to them”. Green ICT is dealing with the environmental impact of the ICT sector itself. ICT for Green reveals how ICT can be applied to make other sectors greener. The ICT sector gives about 2% of the world’s GHG emissions, mainly due to the emissions from the aviation segment. Although, this figure does not look so much, some authors consider that ICT sector emissions are the highest growing one with rising rates of 6% annually. Besides, the environmental impact of ICT is largely neglected, aviation begins to pay attention to the environment decades ago, but the ICT sector has only started to take care of the environment nowadays. Concerning the Green ICT, the main ICT used are presented in Fig. 2. For each of these categories of equipment, there are different options, which can be made to reduce the environmental impact. When discussing the environmental impact of ICT, the most debated topic is energy efficiency. Meanwhile, there is also a question about such estimation of materials used for production, and the way to be done.
Measuring the energy efficiency is the easiest metric to estimate the effect of green ICT, because of the simple way to determine the energy consumption. More energy can at times be used during the equipment production in comparison to its entire lifespan and this is the case for PCs. When producing computers, several metals like aluminum, arsenic, copper, and lead are used. Some of them are hazardous and create needs for handling the equipment, especially during the recycling. It is claimed that 70% of all hazardous wastes are e-wastes. There are regulations for e-waste recycling in a lot of countries, but informal recycling offers a cheap way to deal with this waste and unfortunately is practiced a lot.

Figure 2. Main categories of ICT used.
- Going green in ICT
The term ‘green’ is used in everyday language to refer to environmentally sustainable activities. Also, the term ‘green’ repeatedly is used as being sustainable. Both concepts are closely linked but are not identical and to be green is only a part of being sustainable. The United Nations commonly discuss three aspects of sustainable development: the social, economic, and environmental dimensions. The term ‘green’ fits the environmental aspect because both are often interrelated. From an environmental point of view, it is well understood and reasonable especially in the case of linking ‘green’ ICT with economic growth. Thus, there is a reason that sustainability becomes a factor for the rapid growth of the ICT sector; it favors economic growth, ensures wide access to new technologies, and enhances the efficiency of other sectors.
The sustainability or its environmental issues definition given in the UN report “Our Common Future” said that sustainable are those processes that “meet the needs of the present without compromising the ability of future generations to meet their own needs“.
To be ‘green’ in the ICT sector is partly about taking informed solutions for the way of the use of natural resources.
A good illustration of the term “sustainability “is the way of how today’s living can impact the lives of future generations in terms of lack of some natural resources, health problems for those working with production and recycling, as well as the impact on the future generations. So, dealing with ICT there are different ways to become green.
It is necessary to define ‘green’ technologies and ‘green’ behavior. Green technologies involve parts like virtual servers, allowing a higher rate of utilization to realize a saving of energy of up to 85%, compared to standard PCs. However, the technology that is used inefficiently will not be green. In this way, the employees can have a big importance on the environmental trace. Turning off the equipment after leaving the office, and using technology to make green other parts of life, like hosting video conferences instead of traveling, or sharing different equipment like printers between departments, are examples of green bearing in the workplace.
Being ‘green’ in the ICT sector and information management is not precisely the same as being ‘green’ in other sectors. Each sector possesses its environmental issues and mitigation strategies. The ICT sector is peculiar, as causing greening of other sectors and it is not sure how the large contribution can be obtained for reduction of the emissions and achieving a better environment. Nevertheless, this environmental impact of ICT should not be overestimated. To be ‘green’ in the ICT sector is necessary to become aware of production way and the use and recycling the ICT equipment to be obligatorily sustainable. One of the most important aspects of ‘green’ ICT are those concerned with conserving energy, and other resources, such as paper and rare elements. Rare and hazardous materials are used in the production of ICT equipment. Hazardous materials offend not only the environment but also the people working with the manufacture of ICT equipment and its recycling. All these things are linked with ICT and information storage having the biggest impact on the environment. If all equipment is used optimally, it is reasonable to diminish the impact on the environment as well as possibly to save finances.
- The Importance of being ‘green’
Nowadays to be ‘green’ is significant as a business strategy, which is due partly to consumer understanding of environmental protection. Environmental management is a part of the strategic approach for several successful companies, to involve a new concept of ‘green’ management to answer this strategy. Besides, investing in ‘green’ innovations and environmental protection is beneficial to companies from a financial point of view. Being ‘green’ could increase a company’s competitive advantage, as well as bring new market opportunities, and thus make ‘green’ companies more profitable.
At the same time, the dangers of the so-called ‘green-washing’: “the focus must move away from an emphasis on image to an emphasis on substance” are a serious warning. Thus, ‘green-washing’ is the term assigned to companies that try to get the face of being ‘green’, but apply cosmetic measures, rather than actual changes in the way the organization operates. So, if organizations want to make a real ‘green’ change, it can get a positive effect on society. The greening of organizations can bring new investments to improve the environmental situation, creating jobs and wealth. It is considered that investments in ‘green’ ICT can give short-term economic support.
- Benefits of going green
A major chance to support greening is cost scanting. Observing the ‘green’ information technology, it becomes clear that data centers and servers dispose a large amount of energy for keeping run and cooing down. ICT resources cannot be used in their full capacity (a utilization rate is about 12.5%). This figure indicates lack of efficiency in the ICT sector. It is found that 86% of the needs can be arranged by 26% of the current energy utilization and this indicates a big room for cost improvement. Also, another part of ICT can propose considerable saving opportunities in usage in an environmentally friendly way. When turning off printers, as well as other energy-consuming equipment, whenever they are not exploited, resources can be saved and the environmental impact can be reduced.
Besides these options, it is acknowledged that moving towards a ‘green’ knowledge society will require structural changes, suggesting that governments should help to enable this process, and doing so benefits for the society as a whole will be realized.
No conventional definition of ‘green’ ICT exists. ‘Green’ topics could be approached from the perspective of a ‘problem’ (focusing on reducing the emission) or from the perspective of a ‘solution’ (focusing on new green solutions). This is reflected in the approach to ‘green’ ICT. This approach is presented in Fig. 3. It encompasses two segments: greening of ICT and greening with ICT.

Figure 3. The approach to ‘green’ ICT
Greening of ICT in a narrower sense refers to ICTs with low environmental burdens, but using ICT as an enabler reduces environmental impact across the economy outside of the ICT sector. Green ICT, as greening with ICT, is a new concept and even the leading countries and stakeholders have only about few years of experience. It is important to note that both the traditional ‘problem’ approach and the new ‘solution’ approach are needed. Pollution needs to be regulated and companies need incentives to address their emissions. However, ensuring that the new generation of solution providers get the right incentives is equally important.
Reduction of energy consumption and gas emission
ICT can contribute to the reduction of energy consumption and gas emissions through:
- Inventing innovative energy saver systems, technologies and ‘smart’ devices, using ‘smart energy management’;
- Applying energy saver policies using renewable sources, solar energy and photovoltaic, wind energy, bio-fuel, bio-climatic technology, anti-pollutants technology, etc.
- Recycling and reducing e-waste such as old IT systems, chips, PC, hardware, printers, mobile phones, etc.
About 40% of the total energy consumption is due to households. That is why, innovative ‘smart houses’, constructed by green materials, and green architecture exploiting innovative energy sensors are needed. In this context, ICT systems can achieve to measure, manage and reduce electricity consumption and air-conditioning requirements. During the last decade, technical and industrial product manufacturers were essentially obliged to change the direction of their energy consumption, as a result of the economic crisis in addition to the increased environmental awareness of the public. The concern is taken by the producers towards energy reduction via every computing device, from the laptops and mobiles to the data centers, and will be presumably successful. Consumers show their preference for smart devices, new less energy-consuming technologies, renewable energy sources, and updated, more efficient cooling systems with improved energy management software. These products are equipped with official certification to meet or exceed efficiency guidelines.
- Ways of Greening ICT
It is possible to use ICT in a way that reduces stress on the environment in comparison to the traditional ways. This can be done through new technologies, techniques, and strategies that allow the consumption of less energy and resources. An important part of the ICTs job is to save information. The need for storage is rapidly increasing because of the growth of Internet usage, new laws and regulations, arranging the rules for keeping the information, and scientific computing. The final part of the stored information is collected in data centers. These data centers use enough energy to partially shift the positive effects that ICT can have on society, for which it is claimed that “a fraction of energy savings in ICT and networks could lead to significant financial and carbon savings”. To reach energy efficiency in information storage, it is necessary to choose hardware with better energy efficiency or saving energy methods for the equipment use, which are suggested below.
During the last few years, there is a tendency of awareness rising regarding the impact of modern societies on the environment. A lot of environmentally important factors like energy consumption or e-waste are caused by the application of ICT but at the same time, this technology also could solve some other environmental issues. The dual nature of the issue is approached from an integrative perspective with the creation of the Green Computing concept. It is introduced in 1992 by the United States Environmental Protection Agency starting the ‘Energy Star’ programme. In the beginning, it was enforced on various products such as computer monitors, television sets, and air conditioners. The first well-known result of green computing was the sleep mode option of computer monitors consuming low energy in case user activity is lacking after a certain period. Today green computing gathers a lot of other concepts like confirming computer hardware using, virtualization software, cloud computing or, magnify the energy efficiency of data centers. Green Computing gathers technologies that can contribute to both decreasing the environmental impact of ICT (‘Green IT’ – greening of ICT) and applying information systems (IS) diminishing the environmental impact of ICT consumables (‘Green IS’ – greening by ICT). This integrative vision combines two complementary approaches, presented in Fig. 4.

Figure 4. The integrative vision for Green Computing
The practical implementation of the Green IT and Green IS can be clarified through the following conceptual value models.
Green IT: Value Model
The Green IT value model can assist to reach the goal of environmental sustainability. This conceptual model is grounded on four elements, as shown in Fig. 5.

Figure 5. The Green IT value model
Green IS value Model
The Green IS value Model is important for Green IS acceptance and its impact on the firm’s environmental enforcement. While companies are under permanent suspense from different regulators, clients, and competitors, some of them are ready to cancel efficiency and effectiveness for environmental issues. Green IS investments to fulfill sustainable business experiences are intended to raise turnover and/or income.
Green IS adoption by a company staff (senior managers) is investigated and explained through a model that describes perception as determined by three basic factors. Three types of strategic initiatives have to be taken into account for IS adoption. All of them are depicted in Fig. 6.

Figure 6. The Green IS value Model
This model could be put into action through a strategy for application of variety of technologies and techniques, like video and teleconferencing, emission management systems, etc.
A framework for management and application of Green IT and Green IS can be established that encompass variety of technologies offering opportunities for reducing the negative environmental impacts in activities by producers and consumers that use ICT. Among these Green Technologies are the following.
Cloud computing
A lot of organizations apply a new computing paradigm – cloud computing, to optimize usage and minimize the cost of their computer servers. In this way, previous costs linked with setting up the IT infrastructures are omitted. Cloud computing allows linking shared infrastructure and balancing IT resources for computing tasks in real-time while reducing gas emissions and maintaining the levels of service. Cloud computing relies on sharing of hardware and software resources that are shared by multiple users and dynamically reallocated per demand. This ‘green’ technology removes ultimately the need for a company to have an on-premise data center, which has a positive effect on the natural environment and its resources. To date, many vendors provide Green IS-focused cloud services and are experiencing significant financial growth rates per annum.
Computer power management
To save energy in computing, a common action is turning off the equipment in case it is not in operation. There is a so-called open industry standard “Advanced Configuration and Power Interface” (ACPI 2) that represents an advanced Green IT practice. It allows direct power control of the computer operating system and its underlying hardware. This standard automatically rules out components like monitors and hard drives when the inactive period of the equipment takes place. The computer power savings device includes different sleep modes of the monitor, hard disk, system standby and hibernation, and different CPU power states.
Data Center Design
Data centers consume a big percentage of energy, more than 100 times than standard commercial buildings. Thus, the efficient energy recycling design of data centers, like recycling of waste heat, gives a considerable positive impact on energy safety and can be used in the following areas:
- IT systems: increase hardware usage (through virtualization, see below); allow computer power savings modes; buy energy effective tools (e.g., computer power supplies, computer processors, solid-state storage devices, terminal servers);
- Main power systems: power management parts (hardware and software devices for optimizing work and power); rising the usage of renewable energy; use of natural light instead of electricity;
- Cooling systems;
- Air management.
IT Virtualization
The IT virtualization refers to the separation of IT resources through server virtualization and storage virtualization. The server virtualization means operation of many logical computer systems on one physical hardware, and the storage virtualization means summing physical storage from many network memory-tools, on which to be put a single storage device managed from a central console. The IT virtualization decreases costs for hardware, diminishes energy consumption and physical space use. At the same time, it refines software testing and spreading and rises the versatility of hardware investments. Also, it can help in work-sharing: the servers are either busy or put in a low-power sleep state. Thus, virtualization is one of the main approaches for organizations to implement environmental sustainability into IT practices and business.
Material recycling and e-waste
Exploring proper recycling of IT techniques, e-waste is reduced and harmful elements like lead, mercury, and cadmium are discarded from the natural environment. Then it they can be reused and de novo production might be omitted. The process of materials recycling is feasible nearly for any computing device and accessory, such as hard disks, printer cartridges, and batteries. In this way, persons and business organizations can extend the life of IT equipment by upgrading tools instead of just changing them.
Smart grid and smart matters
The Smart Grid technology encompasses hardware and software that ensures more effective exploitation of actual infrastructures for the generation of electricity, its transmission, and distribution. This is an updated version of the electricity delivery system, automatically acting on information flows. The task is to make better the efficacy, safety, and sustainability of the yielding and spreading of electricity by load balancing and peak load management. Devices on the network possess sensors to collect data (power meters, voltage sensors, fault detectors, etc.). Such a device is the Smart Meter, that makes use of digital communication between devices connected to the grid. The collected data relates to energy production and consuming behaviors of both suppliers and consumers. In that sense, real-time information exchange between producers and consumers helps to better control energy demands and reduces the need for energy surplus during peak hours. Tools on the network have sensors to gather data (power meters, voltage sensors, fault detectors, etc.). They read the energy (electricity, gas, etc.) consumption in defined periods and daily sends data back to the public utility e.g., for monitoring and cost estimation. It can also be used to provide information about energy consumption and to set real-time energy prices to consumers. A key feature is automation technology that lets the public utility adjust and control each device from a central location. Benefits include handling alternative sources of electricity (e.g., solar and wind power), smart control for eco-friendly buildings, and in due course integrating electric vehicles onto the grid.
The future ICT trends: from green to wise
Within the ICT industry, there are two main segments:
- Telecom infrastructure, i.e., telecom networks comprised of base stations and access points that provide connectivity to devices (human-centered or machine-centered);
- Mobile devices/terminals, that are communicated by connecting to the infrastructure (e.g., mobile phones, tablets, sensors, actuators, vehicles, drones, etc.).
In the future, billions of devices/terminals will be connected to millions of base stations. Although both segments (network infrastructure as well as devices/terminals) would benefit from renewable energy, the role of renewable energy is much greater in the network segment due to the following reasons:
- Energy consumption of base stations (telecom infrastructure) is much greater than the energy consumption of devices/terminals;
- Global network coverage is important to realize the envisioned networked society and internet of things;
- Base stations have a larger size and are expensive. It is affordable to integrate a renewable energy system with each base station.
- Base stations need reliable and continuous power provision unlike devices (e.g., mobile phones) that can be charged whenever power is available.
To comprehend the future green ICT trends, it is important to understand not only the history of green but also the history of ICT. For the last 20–30 years the ICT infrastructure has been built, the performance and density will continue to improve and increase. However, a turning point has now been reached. The current threshold is similar to the turning point all industrial eras have experienced. During the installation phase, new solutions are used to increase the efficiency in the old system, during the deployment phase the new system reaches maturity, allowing it to deliver entirely new solutions. Initially, the transformation happened in the “information sectors”, within e.g., music, video or book sectors, etc., and now we start to see the first signs of a serious change in the “heavy sectors”, such as car and mobility, construction, agriculture, and retail sectors, as well as in basic business models.
The shift from improving existing systems to providing new solutions is supported by two trends of ICT development and ‘green’ ICT that are important to understand.
- (1) the ICT companies are now influential economic players. For the first time in history, an ICT company – Apple – was the largest company in the US. Apple overtook Exxon Mobile, demonstrating that ICT companies can no longer be ignored by policymakers.
- (2) the ICT companies are now part of a ubiquitous network that is connecting almost everyone and almost everything on the planet. Nowadays, more people are connected than during any other time in human history. By 2020, there are about 50 billion connected devices, and the society is gaining access to data and experience transparency that is fundamentally different from what any society has ever experienced before.
One of the major challenges is that the new ‘green’ ICT solutions have to compete in a regulatory environment encompassing regulations created for the 19th-century industrial structure. It has also to deal with the predispositions among people unfamiliar with the speedy development of ICT solutions. In fact, the current technological development is so fast that society members of any kind – from policymakers to business leaders, to economic experts are witnessing how the whole procedure – from an idea to full-scale implementation, is taking place just for few decades. To clarify the situation, the phases of a long-term disruptive solution must be studied. At the beginning, something triggers an idea that spreads, e.g., the first personal computers. These brought about next ideas, such as paperless offices, virtual meetings, etc. All these activities excited people. After a while, working prototypes were introduced and many companies invested in very expensive prototypes of videoconference equipment. The technology was too new and no viable business model was used, instead, these prototypes were bough and managed by the companies themselves. Progressively, many decision-makers had been excited by the idea of ICT as a disruptive force in different areas, that they finally thought it would never come to happen.
A trend today is that policymakers do not consider significant changes that will be made by ‘green’ ICT. Many are still planning to invest in new coal power plants because they use the same economic model as they used previously. For instance, Siemens Germany has announced that they support a 25% target for renewables and are ending their nuclear power business.
Conclusions
The history of ICT’s development shows that society now is at an inflection point where ICT solutions move towards creating new solutions instead of making old systems better. Two important trends concerning ICT development could be pointed out:
- (1) The fact that ICT companies now are economically powerful and serve as a source of both economic and political capital.
- (2) ICT solutions and companies now are so abundant that new clusters of solutions providers can emerge. It is important to understand what ICT solutions actually can deliver.
‘Green’ ICT includes the use of ICT solutions to support smart growth. Consequently, to observe ‘green’ ICT, the context of current ‘green’ ICT trends should be comprehended. While until the 21st century, ‘green’ was seen as nature conservation or pollution control, a significant shift took place at the beginning of the 2000s. A focus moved from a problem perspective (pollution control) to a solution perspective and a new generation of business leaders saw the opportunity to link the need for resource efficiency and sales of new products and services.
A parallel shift is observed in the ICT development. It moves from improving existing systems to providing new solutions. New ICT solutions were created to support energy efficiency and ‘green’ growth. One of the most popular examples for ICT-driven solutions are the e-books, smart grids, electric cars, online meetings, etc.
It’s positive that many citizens have begun to realize the concept of human-caused climate change and resource depletion. They consider as well the imperative necessity of acting on this matter. This understanding has led many people to make changes of their personal lifestyle. The “living green” tendency includes many aspects such as green constructions, renewable energy use, energy-saving at home, the extended use of eco-friendly products, a recycling approach. Here to add are the so-called sustainability checklists, designed to help households to assess how sustainable they are, and to offer suggestions for increasing home sustainability. Moreover, clever use of e-services can be a tool for less energy consumption in everyday life and at work. For instance, this is the case of paper use – the nowadays correspondence is substituted by digital formats using the Internet and smart devices. The production and distribution of new products and services show the tendency for minimization of the needed energy, estimated by carbon footprint. Another example is the substitution of traditional conferences with online ones that impose a direct positive effect on environmental protection and reduction of the GHG emissions because of reduced transportation services. It is reported that teleconferences can avoid the production of approximately 540.000 tn CO2 per year; this is the cost of the air transportation of people.
The use of ‘Green’ ICT tools and services through broadband/5G Internet contributes to the environmental and societal welfare with the decrease of cost and time to access government offices (24/7 services), energy savings (no transportation), and restriction of pollute emission (carbon footprint). Similarly, in the sector of e-commerce and e-business new innovative business solutions are in favor of either the entrepreneur or the final customer.
Test: LO4 Basic level
References
- Andreopoulou ZS. 2012. Green Informatics: ICT for Green and Sustainability. Journal of Agricultural Informatics. 3, 2, 1-8.
- Berl A, Gelenbe E, Di Girolamo M, Giu G. 2010. Energy-Efficient Cloud Computing. The Computer Journal 53(7), DOI: 10.1093/comjnl/bxp080
Brush, K, Kirsch B. Virtualization. https://searchservervirtualization.techtarget.com/definition/virtualization 22/03/21 - EU ENERGY STAR programme; https://ec.europa.eu/energy/topics/energy-efficiency/energy-efficient-products/energy-star_en
Environmental Technology. http://en.wikipedia.org/wiki/Environmental_technology 22/03/21 - Gholami R, Sulaiman A, Ramayah T. Molla A. 2014. Senior Managers’ Perception on Green Information Systems (IS) Adoption and Business Value: Results from a Field Survey. Information & Management 50(7):431-438, DOI: 10.1016/j.im.2013.01.004.
- ICT for Sustainable Growth: Energy Efficiency of the ICT Sector. DAE Actions. https://ec.europa.eu/information_society/activities/ sustainable_growth/ict_sector/index_en.htm 22/03/21
- Joumaa C, Kadry S. 2012. Green IT: Case Studies. Energy Procedia, 16, 1052 – 1058.
- Klimova S. 2016. Systematic literature review of using knowledge management systems and processes in green ICT and ICT for greening. Conference: International SEEDS Conference, Leeds Beckett University, UK, 1-21.
- Malmodin J, Lundén D. 2018. The Energy and Carbon Footprint of the Global ICT and E&M Sectors 2010–2015. Sustainability, 10, 3027; doi:10.3390/su10093027
- OECD Towards Green ICT Strategies: Assessing Policies and Programs on ICT and the Environment. http://www.oecd.org/dataoecd/47/12/42825130.pdf 22/03/21
- OECD countries agree to tackle global environmental challenges through information and communication technologies (ICTs). http://www.oecd.org/document/26/0,3343,en_2649_33757_45073498_1_1_1_1,00.html 22/03/21
- OECD. Towards Green ICT Strategies: Assessing Policies and Programmes on ICT and the Environment. http://www.oecd.org/dataoecd/47/12/42825130.pdf 22/03/21
- Porter ME, Kramer MR. 2006. Strategy and Society: The Link Between Competitive Advantage and Corporate Social Responsibility. Harvard Business Review, https://hbr.org/2006/12/strategy-and-society-the-link-between-competitive-advantage-and-corporate-social-responsibility
- Report of the World Commission on Environment and Development: Our Common Future. Brundtland, 1987
- UN ECONOMIC and SOCIAL COUNCIL. Sustainable Development https://www.un.org/ecosoc/en/sustainable-development
- UN Report “Our Common Future” https://sustainabledevelopment.un.org/content/documents/5987our-common-future.pdf
- White paper: GREEN COMPUTING. 2016. https://portail-qualite.public.lu/dam-assets/fr/publications/normes-normalisation/information-sensibilisation/white-paper-green-computing/white-paper-green-computing.pdf 22/03/21
- Williams G, Duncan A. Landell‐Mills P, Unsworth S. 2010. Politics and Growth. Dev. Policy Rev., 5-31, https://doi.org/10.1111/j.1467-7679.2011.00519.x
- Yadav K. 2014. Green Computing – A Necessity Now. https://www.acecloudhosting.com/blog/green-computing-a-necessity-now/ 22/03/21
- Andreopoulou ZS. 2012. Green Informatics: ICT for Green and Sustainability. Journal of Agricultural Informatics. 3, 2, 1-8.
- Berl A, Gelenbe E, Di Girolamo M, Giu G. 2010. Energy-Efficient Cloud Computing. The Computer Journal 53(7), DOI: 10.1093/comjnl/bxp080
Brush, K, Kirsch B. Virtualization. https://searchservervirtualization.techtarget.com/definition/virtualization 22/03/21 - EU ENERGY STAR programme; https://ec.europa.eu/energy/topics/energy-efficiency/energy-efficient-products/energy-star_en
Environmental Technology. http://en.wikipedia.org/wiki/Environmental_technology 22/03/21 - Gholami R, Sulaiman A, Ramayah T. Molla A. 2014. Senior Managers’ Perception on Green Information Systems (IS) Adoption and Business Value: Results from a Field Survey. Information & Management 50(7):431-438, DOI: 10.1016/j.im.2013.01.004.
- ICT for Sustainable Growth: Energy Efficiency of the ICT Sector. DAE Actions. https://ec.europa.eu/information_society/activities/ sustainable_growth/ict_sector/index_en.htm 22/03/21
- Joumaa C, Kadry S. 2012. Green IT: Case Studies. Energy Procedia, 16, 1052 – 1058.
- Klimova S. 2016. Systematic literature review of using knowledge management systems and processes in green ICT and ICT for greening. Conference: International SEEDS Conference, Leeds Beckett University, UK, 1-21.
- Malmodin J, Lundén D. 2018. The Energy and Carbon Footprint of the Global ICT and E&M Sectors 2010–2015. Sustainability, 10, 3027; doi:10.3390/su10093027
- OECD Towards Green ICT Strategies: Assessing Policies and Programs on ICT and the Environment. http://www.oecd.org/dataoecd/47/12/42825130.pdf 22/03/21
- OECD countries agree to tackle global environmental challenges through information and communication technologies (ICTs). http://www.oecd.org/document/26/0,3343,en_2649_33757_45073498_1_1_1_1,00.html 22/03/21
- OECD. Towards Green ICT Strategies: Assessing Policies and Programmes on ICT and the Environment. http://www.oecd.org/dataoecd/47/12/42825130.pdf 22/03/21
- Porter ME, Kramer MR. 2006. Strategy and Society: The Link Between Competitive Advantage and Corporate Social Responsibility. Harvard Business Review, https://hbr.org/2006/12/strategy-and-society-the-link-between-competitive-advantage-and-corporate-social-responsibility
- Report of the World Commission on Environment and Development: Our Common Future. Brundtland, 1987
- UN ECONOMIC and SOCIAL COUNCIL. Sustainable Development https://www.un.org/ecosoc/en/sustainable-development
- UN Report “Our Common Future” https://sustainabledevelopment.un.org/content/documents/5987our-common-future.pdf
- White paper: GREEN COMPUTING. 2016. https://portail-qualite.public.lu/dam-assets/fr/publications/normes-normalisation/information-sensibilisation/white-paper-green-computing/white-paper-green-computing.pdf 22/03/21
- Williams G, Duncan A. Landell‐Mills P, Unsworth S. 2010. Politics and Growth. Dev. Policy Rev., 5-31, https://doi.org/10.1111/j.1467-7679.2011.00519.x
- Yadav K. 2014. Green Computing – A Necessity Now. https://www.acecloudhosting.com/blog/green-computing-a-necessity-now/ 22/03/21
Open access scientific resources: digital databases
B A S I C L E V E L
An Open access or OA is a set of principles and a range of practices through which research outputs are distributed online, free of cost or other access barriers.
Open access scientific resources
An introduction to « Open Access » resources
An Open access or OA is a set of principles and a range of practices through which research outputs are distributed online, free of cost or other access barriers. With open access strictly defined (according to the 2001 definition), or libre open access, barriers to copying or reuse are also reduced or removed by applying an open license for copyright.
The main focus of the open access movement is “peer reviewed research literature”. Historically, this has centered mainly on print-based academic journals. Whereas conventional (non-open access) journals cover publishing costs through access tolls, such as subscriptions, site licenses or pay-per-view charges, open-access journals are characterized by funding models which do not require the reader to pay to read the journal’s contents. Open access can be applied to all forms of published research output, including peer-reviewed and non-peer-reviewed academic journal articles, conference papers, theses, book chapters, monographs and images.
However, when it comes to define “free” access, one has to distinguish “gratis” from “libre”.
In order to reflect real-world differences in the degree of open access, the distinction between gratis open access and libre open access was added in 2006 by Peter Suber and Stevan Harnad, two of the co-drafters of the original Budapest Open Access Initiative (BOAI) definition of open access publishing. Gratis open access refers to online access free of charge and libre open access refers to online access free of charge plus some additional re-use rights. Libre open access is equivalent to the definition of open access in the BOAI, the Bethesda Statement on Open Access Publishing and the Berlin Declaration on Open Access to Knowledge in the Sciences and Humanities. The re-use rights of libre OA are often specified by various specific Creative Commons licenses; these almost all require attribution of authorship to the original authors.
The document released in February 2002 by the BOAI contains the following very widely used definition:
- By “open access” to this literature, we mean its free availability on the public internet, permitting any users to read, download, copy, distribute, print, search, or link to the full texts of these articles, crawl them for indexing, pass them as data to software, or use them for any other lawful purpose, without financial, legal or technical barriers other than those inseparable from gaining access to the internet itself. The only constraint on reproduction and distribution and the only role for copyright in this domain, should be to give authors control over the integrity of their work and the right to be properly acknowledged and cited.
In light of the information above, the use of open source scientific resources must follow the rules commonly adopted. Publishing open source scientific resources must also clearly mention if they are libre or gratis and must be attributed to the original author.
An introduction to data (basic level)
What are “data”
According to the Merriam-Webster dictionary, there are three different definitions of data:
- Factual information, such as measurements or statistics, used as a basis for reasoning, discussion, or calculation
- Information in digital form that can be transmitted or processed
- Information output by a sensing device or organ that includes both useful and irrelevant or redundant information and must be processed to be meaningful
In this document we will cover most of the three definitions.
A brief history of data
Since humans began to communicate, they have experienced the need to retain information for the long term. Keeping information was necessary for our ancestors to ensure their survival. Transmitting information across generations allowed them to keep track of potential dangers, but also to have an inventory of the best places to collect food, the best spots for fishing, the most interesting animals to hunt and where to find the best shelters. All of this information was transmitted orally. With the evolution of knowledge and the invention of writing, they began to store information on indelible media.
Without going into detail about the evolution of the representation of information, some significant examples will be provided that have helped in the structuring of thought, which lead to the discovery of the computer tools we use daily.
Data prior to the invention of computers
As human societies emerged, collective motivations for the development of writing were driven by pragmatic exigencies. These include organizing and governing societies through the formation of legal systems, contracts, deeds of ownership, taxation, trade agreements, treaties, census records, keeping history, maintaining culture, keeping track of scientific discoveries, codifying knowledge through curricula and lists of texts that are artistically exceptional or deemed to contain

Figure 1: Cuneiform writing
foundational knowledge, and many other needs.
For example, around the 4th millennium BC, the complexity of trade and administration in Mesopotamia outgrew human memory, and writing became a more dependable method of recording and presenting transactions in a permanent form.
Cuneiform was one of the earliest systems of writing, invented by Sumerians in ancient Mesopotamia. It is distinguished by its wedge-shaped marks on clay tablets, made by means of a blunt reed for a stylus, as demonstrated in Fig. 1.
Over time, the development of knowledge, the multiplication of information, the limitation of human memory, the necessity of writing and keeping record of huge quantities of information has become essential. However, despite keeping record of almost every kind of information or data on various media, it became more and more complex to retrieve it in a simple manner. One had to read tens of reports and books to be able to synthesize on a subject
Data in modern age
Today, the quantity of data produced every year and kept digitally, e.g., to-do lists, recipes, reminders, logbooks, maps, photos, e-mails, scientific data, political reports, videos, etc. is so exponential that it creates the need to structure the way we can retrieve these phenomenal quantities.
Computers gained popularity and became cost effective to use by individuals and private companies in the early 80’s. However, the 60’s can be considered as the new era in the field of databases. The introduction of the term “database” coincided with the availability of direct-access storage or DAS, from the mid-60s onward. This new technology represented a contrast with the past punch cards and the tape-based systems, allowing shared interactive use rather than daily batch processing. Two main data models were developed – network model “CODASYL” (Conference on Data System Language) and hierarchical model “IMS” (Information Management System).
The first generation of database systems was “navigational, in opposition to the sequential access due to the previous technologies used to store data, i.e. tapes and punch cards. Applications typically accessed data by following pointers from one record to another. Storage details depended on the type of data to be stored.
Adding an extra field to a database required rewriting the underlying access/modification scheme. Emphasis was on records to be processed, not the overall structure of the system. A user would need to know the physical structure of the database in order to query for information. One database that proved to be a commercial success was the “SABRE” system that was used by IBM to help American Airlines manage its reservations data. This system is still utilized by the major travel services for their reservation systems.
In modern Information technology, confusion has always existed among users between databases and internet web search engines accessed by browsers. A database usually contains structured data, in contrast to the World Wide Web (www), which usually contains unstructured data. Even if retrieving information from both databases and “www” are seamless and look similar, the content and the way queries are addressed are completely different. Structured and unstructured data will be explained later in this document.
Understanding the basic vocabulary
Terminology
Like any other science, computer science has its own language. In order to fully comprehend the information that will be provided in this document, it is essential to become familiar with the vocabulary related to this topic.
Moreover, the communication with a DBA (Database Administrator) will be eased. When a Biochemist will have to express his needs in terms of structuring or managing data in a Database, he will be tempted to use his own technical language. Then the DBA will have to understand the request and transform it into a computer language, which will be understandable by biochemists.
What are “Data” in the computer age

As mentioned in section 2.1, according to the domain that is being referred to, data might have different meanings. In the case of computing and databases, data is defined as any sequence of one or more symbols. Data requires interpretation to become information. In information technology the “bit” is the smallest quantity of data. A bit is binary. Binary numbers are a representation of numbers using only two digits, 0 and 1 (Fig. 2). It is a base-2 numeral system, i.e.:
- 0 0 0 1 = numerical value 20
- 0 0 1 0 = numerical value 21
- 0 1 0 0 = numerical value 22
- 1 0 0 0 = numerical value 23

A sequence of “bits” constitutes a “Byte”. Bytes are made of a multiple of 4 bits (a byte of 4 bits is called a Nibble) as in the example above. Today, the byte is a unit of digital information that most commonly consists of eight bits. Historically, the byte was the number of bits used to encode a single character of text in a computer. With a byte of eight bits the maximum decimal number is 256. Historically, the byte is also the unit of computer information or data-storage capacity used to measure the quantity of data (Table 1).

An example of usage is the ASCII (American Standard Code for Information Interchange) table of characters commonly used for alphabetical characters (Table 2). The first 32 characters are called control characters. Initially, they were not designed to represent printable information, but to control devices that use ASCII code, such as printers, or to provide meta-information about data streams, e.g., those stored on magnetic tape.
What is “Metadata”

Metadata, or, put simply, meta-information, is used to reference the data about the data. Having data is not enough to simply put them online. Data are not usable until they can be explained in a manner that both humans and computers can process.
Metadata may be implied, specified or given. It includes data relating to physical events, or processes, and will also have a temporal component. In almost all cases this temporal component is implied. It may be slightly tricky to understand, however, the following example will provide a clearer explanation of this term.

Metadata of the picture
Imagine that you are traveling with your favorite smartphone in some paradisiac island. You start taking pictures (Fig. 3) to keep nice records of your trip. A week later, your trip reaches its end and you have to go back home.
Back home, you invite your best friends for a party and want to share with them the beauties you have seen during your trip. You start showing the pictures, but you cannot recall which day, at what time and where some of them were taken. This is where the metadata of the pictures can help. In a few words, it is the description of the data. In this example, the picture is the data and the description of the picture is the metadata (Fig. 4).
In Biotechnology, one must understand that metadata are by far more important than data. It is very simple to understand the reason why metadata are a crucial component directly related to data. Imagine an experiment that will lead to a specific result. This experiment, to be valid, must be documented. This documentation should include all the conditions, under which the experiment was conducted. This might include the description of the kind of raw material used, its source, in which conditions it was collected, the types of machines to process the experiment, temperature, date, time, etc. For the result of this experiment to be comparable to other results of similar experiments, all the conditions must be similar. Raw data without metadata are useless.
The biggest challenge in Biotech, and any other science, is to standardize metadata. In most of the Biotech databases, this is not respected. One must absolutely be conscious of this phenomenon and thoroughly respect the standards.
What is a “Database”

In general, a database is defined as a collection of data items, such as phone books, price lists, inventory lists, customer’s addresses, etc. Nonetheless, in technical terms, a database is referred to as “a self-describing collection of integrated records”. It implies computer technology, completed with a specific computer language, such as SQL (Structured Query Language).
A database consists of multiple tables (Fig. 5) and of both data and metadata. Metadata is the data that describes the structure of the data within a database. If you know how your data is arranged, then you can retrieve it. Since the database contains a description of its own structure, it is referred to as self-describing. The database is integrated because it includes not only data items but also the relationships among them.
The database stores metadata in an area called the data dictionary, which describes the tables, columns, indexes, constraints and other items that make up the database.
Because a flat file system i.e. “Spreadsheet” has no metadata, applications written to work with flat files must contain the equivalent of the metadata as part of the application program.
What are “Tables” in a database
A table is a collection of related data held in a table format composed of columns and rows within a database. It resembles a spreadsheet (Fig. 6).
What are “Columns” in a database

A column is a set of data values, all of a single type, in a table. Columns define the data in a table. Most databases allow columns to contain complex data like images, whole documents or even video clips. Therefore, a column allowing data values of a single type does not necessarily mean it only has simple text values. Some databases go even further and allow the data to be stored as a file on the operating system, while the column data only contains a pointer or link to the actual file. This is done for the purpose of keeping the overall database size manageable – a smaller database size means less time taken for backups and less time required to search for data within the database.
In a table, each column is typically assigned a data type and other constraints, which determine the type of value that can be stored in that column. For example, one column might accept email addresses, and another might accept phone numbers with a constraint of 10 digits.
What is a “Record”
A record is a representation of a physical or conceptual object. Say, for example, that you want to keep track of the customers of a business. You assign a record for each customer. Each record has multiple attributes, such as name, address, and telephone number. Individual names, addresses and so on are the data.
What are “Indexes”

Structured data are stored in the form of records in a database. Every record has a key field, which helps it to be recognized uniquely, i.e. the ID of a patient. No other patient can have the same ID number, but another patient may have the same first name and last name.
Indexing a database is a technique to efficiently retrieve records from the database files, based on some attributes on which the indexing has been performed. To make it simple, indexing in database systems is similar to what we usually see in books. At the beginning or the end of a book, an index may be found (which is different from a table of contents), which provides all the page numbers for a specific topic. For example, an Atlas may be divided into chapters containing maps, chapters containing data on population and chapters dedicated to countries production or agricultural data. If you are looking for a specific country and you would like to have an overview of all the data regarding this specific country, the index might be very helpful as it will show you the page related to that country in each chapter (Fig. 7).
What is an “Object”
In computer science, an object can be a variable, a data structure, a function, or a method and, as such, is a value in memory referenced by an identifier. In the relational model of database management, an object can be a table or column, or an association between data and a database entity, such as relating a person’s age to a specific person.
Structured data

According to SNIA (Storage Networking Industry association), structured data is defined as:
“Data that is organized and formatted in a known and fixed way.
The format and organization are customarily defined in a schema. The term structured data is usually taken to mean data generated and maintained by databases and business applications.”
Three conditions are needed to describe data as structured:
- It must conform to a data model,
- It must have a well define structure,
- It must follow a consistent order and can be easily accessed and used by a person or a computer program.
Structured data is usually stored in well-defined schemas such as Databases. It is generally tabular with columns and rows that clearly define its attributes (Fig. 8).
SQL (Structured Query language) is often used to manage structured data stored in databases.
Unstructured data

Information that is not organized in a predefined model is called unstructured data or unstructured information. In computer science, files like text files, photos, video files, audio files and presentations are considered unstructured files. Typically, a PDF file contains unstructured data (Fig. 9).
It is estimated that 80 to 90% of the worldwide total dematerialized data is unstructured. Usual query algorithms are unable to simply and efficiently extract the required information from an unstructured file, such as in the example of Fig. 9. The same information contained in Fig. 9 can easily be retrieved with a query. However, today, unstructured data analytics tools powered by artificial intelligence (AI) are available, which were specifically created to access the insights available from unstructured data (see 3.1.12 Analytics).
Big data
According to SNIA (Storage Networking Industry association), big data is defined as:
“A characterization of datasets that are too large to be efficiently processed in their entirety by the most powerful standard computational platforms available.”
In other words, Big Data refers to huge quantities of structured or unstructured data that cannot be processed by usual software as traditional database query language or any other kind of fetching engine.
Confusion exists concerning the current usage of the terms Big Data and Analytics. Big Data is the information, while Analytics is the way to extract the desired information from huge quantities of available information.
Analytics
In computer technology, Analytics is a method to extract value from big data.
In the field of healthcare, Big Data Analytics has led to many improvements by providing personalized medicine and predictive analytics. As the volume of data is dramatically increasing, traditional databases and search engines will not be able to handle and retrieve specific information. Patient data is generated by MRI’s, X-rays, blood tests machines, monitoring sensors and many more sources of data complex to process. Extensive information in healthcare is now in electronic form; it fits under the big data umbrella as most of it is unstructured and difficult to use.
Big data in health research is particularly promising in terms of exploratory biomedical research, as data-driven analysis can move forward more quickly than hypothesis-driven research. Subsequently, trends seen in data analysis can be tested in traditional, hypothesis-driven follow-up biological research and eventually clinical research.
Repository
A data repository or data warehouse is a centralized place to store and maintain data. A data repository can consist of one or more structured data files, such as databases or unstructured data files, which can be distributed over a network and preserved over the long-term.
Basic structure of a database
This section is dedicated to the overview of the main building blocks constitutive of a database.
Introduction
Since the invention of computers, the amount of data stored and managed electronically has increased drastically. It is estimated that the quantity of data will reach 175 zettabytes (1021 Bytes) by 2025 growing from a few petabytes (1015 Bytes) in the year 2000. One common way of simplifying the lives of users and making the most of their resources is by storing and retrieving it more efficiently. For example, while a flat file works just fine for storing your personal data, such as an address book or some recipes, it is not as suitable for storing a city phone directory or, more precisely, the genomic data in the Biotech field. In addition, if you want to store several genomic species worth of data, it is very difficult to search and retrieve data from a flat file. Databases offer a solution to this problem by making the storage, handling and retrieval of data much easier.
The software used to manage a database is called a database management system (DBMS). This specialized software acts as a go between to help end users access the database. Usually, users do not interact directly with a database because this may result in its disorganization. Instead, they use a DBMS that reads data from or writes data to the database.
The growing complexity of big quantities of data required some companies to use data management tools based on the relational model, such as the classic RDMBS. RDBMS stands for Relational Database Management System. Nevertheless, major Internet companies, such as Google, Yahoo and Amazon, or all the popular Social Media, each faced challenge in dealing with huge quantities of data in real-time, something that conventional RDBMS solutions could not cope with. That explains the soaring popularity of NoSQL database systems that sprang up alongside.
NoSQL systems are distributed, non-relational databases designed for large-scale data storage and for massively-parallel, high-performance data processing across a large number of commodity servers. They arose out of a need for agility, performance and scale, and can support a wide set of use cases, including exploratory and predictive analytics in real-time. Built by top internet companies to keep pace with the data deluge, NoSQL databases scale horizontally and are designed to scale to hundreds of millions and even billions of users performing updates as well as reads.
Some of the common applications of NoSQL databases are social media, large scale e-mail providers and governmental healthcare systems.
Usually, a social application can scale from zero to millions of users in a few weeks and to better manage this growth, one needs a DB that can manage a massive number of users and data, but can also easily scale horizontally.
In this course, we will focus on DBMS and RDBMS only. These are the two kinds of Databases commonly used in the Biotech world up to date.
Overview of a database architecture

Databases can store all kinds of information, from numbers and text, to email, web content, phone records, biological, geographical data, etc. Databases are officially classified according to how they store this data. Relational databases store data in tables. Object oriented databases store data in object classes and subclasses. We are going to focus on relational databases, as they are most commonly used. However, most of the basic topologies of databases need to have backend servers in order to host the database management system, a storage system attached to the servers to store the structure and the data of the database and, of course, computers, laptops, desktops or terminals as an interface to allow users to access the database, its management system and its content. Also required is a network to exchange between all the hardware components and a Cloud attachment to allow remote users to access the database. Fig. 10 summarizes in a simple way the minimum required to run a database.
Another basic way to describe it, is to show the three level architecture of a database. It is a virtual view of the necessary layers to make a database function properly. Fig. 11 demonstrates the three-level view architecture. It is called the ANSI-SPARC model. Nonetheless, despite the fact that this model never became a formal standard, it presents the idea of logical data independence that has been widely adopted.
Information stored inside a relational database is contained within tables. These tables are composed of rows of data and each row contains fields or columns. In a well-designed database definition, called a schema, only similar data is stored within each table and duplication of columns is kept to a minimum. Developers can connect, or join, data from two tables to link different types of information to each other.

Indexes can be created on fields in the database table to make it easier for the DBMS to retrieve data. Indexes are usually configured for frequently searched columns, like a person’s name or a date value. The drawback to using indexes is that they take up storage disk space and can slow things down, if too many of them are maintained, because every time a row in the database is updated, the index also has to be updated.
Most databases support Structured Query Language (SQL), a standard language for interacting with information contained in a database. SQL allows users and applications to interact with specific subsets of data from one or more tables using several statements as SELECT, INSERT, UPDATE and DELETE.
Relational databases also provide a layered approach to storage, allowing the definition of what database objects reside in specific data files and where those data files are placed within the operating system’s file structure. On top of managing the physical storage location of database objects, many database systems give some control over how the data is stored within the data files.
Common database terms
Certain database terms derive from ways that databases automate write actions. Database developers often automate writing to certain fields or other tables, such as writing a copy of the row being inserted – along with a timestamp or username – to a history or audit table. Most DBMS systems provide several ways to automatically manage database write actions.
Database triggers are the most common method of taking action on data as it is being written to the database. Triggers are usually associated with a particular table and configured to execute at a certain point during a specific write action, such as before or after an update, or after a row is inserted. Triggers can be used to format data, populate a column with data derived from existing information, or even write to another table based on the row being inserted or updated.
A stored procedure is another way of interacting with a relational database. Stored procedures are more complex than triggers and are not tied to a single specific table. Typically created by a developer, they use a combination of SQL and a programming language, such as Java or SQL (depending on the database platform). Stored procedures provide developers a lot of control over how data is validated or massaged by an application. A stored procedure could be used to manage how a user logs in to an application. The procedure might first validate the username and password, then log the success or failure of the attempt to another table, along with other information, including the computer name and a timestamp. An alert could even be sent to the user informing them that their password has expired and must be changed.
Functions are simpler than stored procedure, and can sometimes even be used from within SQL queries. Functions are usually used in a database to perform a set of actions that return one or more values, such as calculating the sum of a column for rows that match a certain condition. While these actions can be performed using SQL, building them into a function can make them easier to use in other code. Both functions and stored procedures can perform common actions in a streamlined and consistent manner, easing the workload for database administrators and developers.
What is the difference between major DBMS systems?
The DBMS is generally driven by what the user applications need to support. That said, here is a brief comparison of the three most widely used platforms.
Microsoft SQL Server is widely used in enterprise applications and integrates easily with other Microsoft tools. Microsoft SQL Server 2019 Express is the latest version of Microsoft’s free offering and is often bundled with applications that use SQL Server.
MySQL has been a favorite for open source developers for the better part of two decades. Often used as a back end for open-source blog or content management systems, MySQL has a massive installed base across the globe. In 2008, MySQL AB was acquired by Sun Microsystems, which was itself acquired by Oracle Corp. in 2009, bringing MySQL under the umbrella of one of its largest competitors. However, the MySQL Community Edition remains free and is well supported by the community. MySQL is available for numerous operating systems, including Linux, UNIX, Mac OS X and Windows.
Oracle Database is considered by many to be the standard in enterprise-level database platforms and supports numerous enterprise applications. Oracle Database Express Edition is available free of charge and is also free to distribute (though it is not technically free software), making it another popular option for developers or hobbyists on Windows or Linux.
Now that you have learned the fundamental database terms and concepts, you are that much closer to speaking the same language as your organization’s database developers.
Databases in the scientific word
This part deals with the basics of databases used in the scientific world
Introduction to existing databases dedicated to science
This section is dedicated to the overview of the most common open access databases used in science.

Continuous developments in the fields of biotechnology and information technology have led to the exponential growth of data. Studies conducted by researchers at the European Bioinformatics Institute (EMBL-EBI) have demonstrated that this growth of information is doubling approximately every year. These extensive amounts of data are stored, organized and constantly updated in scientific databases, where they are readily available for scientists, including biologists and bio-informaticians, to use for research purposes. The information available in biological databases is obtained from a range of scientific fields, including metabolomics, microarray gene expression and proteomics. Apart from storing, organizing and sharing huge volumes of data, the main goal of biological databases is to offer web application programming interfaces (APIs) for computers to exchange and integrate data from many different database resources via an automated method.
Biological databases can be defined as data collections, which are structured in such a way making their contents easy to explore, handle and update. Examples of such databases are presented in Fig. 12. In 1972, the first protein structure database, known as the Protein Data Bank (PDB), was created. This database originally contained only 10 entries, which has now expanded to contain more than 10,000 entries, signifying the rapid growth of biological data. A biological database may contain several types of data, including protein sequences, textual descriptions, attributes, and tabular data. Generally, they can be divided into primary, secondary, and composite databases. Primary databases include data about the sequence or structure alone, whereas secondary databases include data originating from the primary database. Data, such as the conserved sequence and active site residues of protein families, can be found in secondary structure databases. Furthermore, entries of the PDB, which is a primary database, can be found in secondary structure databases, stored in an organized way.
Broadly speaking, biological databases can be categorized into sequence, structure, and pathway databases:
- Sequence databases: The most commonly used biological databases. These include protein and nucleotide sequence databases, which contain wet lab results and are the main source for experimental results. GenBank and EMBL are examples of sequence databases.
- Structure databases: These databases contain information regarding protein structure and molecular interactions. PDB is an example of a structure database.
- Pathway databases: These databases are based on data derived from the comparative study of metabolic pathways. The Kyoto Encyclopedia of Genes and Genomes (KEGG) and Biocyc are two indicative pathway databases.
A typical search in a nucleotide sequence database may, for example, generate data concerning the scientific name of the source organism from which it was isolated, contact name, the input sequence with details of the molecule type and, frequently, literature citations related to the sequence.
Certain tools have been developed to facilitate scientists in data processing and retrieval from biological databases. These tools, which are termed bioinformatics tools, are software programs created for the extraction of meaningful data from the vast number of biological databases and for conducting sequence or structural analysis. Bioinformatics tools are used to obtain data from genomic sequence databases and for the visualization, analysis and retrieval of date from proteomic databases. These tools are largely divided into:
- Homology and similarity tools: These tools are used for the detection of similarities between the sequences of unknown structural and functional sequences, whose function and structure are already known.
- Protein function analysis tools: Programs applied for the comparison of one protein sequence to a secondary (or derived) protein, which permit the estimation of the biochemical function of a query protein.
- Structural analysis tools: These tools allow the comparison of structures with the known structure databases and the establishment of the 2D/3D structure of a protein.
- Sequence analysis tools: Programs used for the additional, more comprehensive assessment of a query sequence, involving evolutionary analysis and identification of mutations.
Biological databases may also be categorized, based on the scope of data coverage, into:
- Comprehensive databases: These databases comprise various types of data from a number of species. Examples of comprehensive databases are GenBank and EMBL.
- Specialized databases: These databases include particular types of data or data from particular organisms. An example of specialized databases is WormBase, which contains information on nematode biology and genomics.
In relation to the level of biocuration, which is defined as the activity of organizing, demonstrating and making biological information readily available to both humans and computers, biological databases are classified as primary and secondary or derivative databases. Primary databases consist of raw data as archival repository, while secondary or derivative databases consist of curated information as added value. In regard to the method employed for curating the data, biological databases may be further classified as expert-curated databases or community-curated databases, which are curated in a co-operative way by numerous researchers.
Additional categorization of biological databases can also be made based on data type. The data types that accordingly classify databases include DNA, RNA, protein, expression, pathway, disease, nomenclature, literature and standard and ontology. Some of the most important and widely used biological databases are the following: GenBank, the UCSC Genome Browser and Ensembl, which are sequence databases/portals; WormBase and The Arabidopsis Information Resource (TAIR), which are model organism databases; and the PDB, Online Mendelian Inheritance in Man (OMIM), MetaCyc and KEGG, which are characterized as non-sequence-centric databases.
Data manipulation is an essential part of the experimental process of all studies, regardless of their scale. The online availability of biological data combined with the decreasing costs of automated genome sequencers have made it possible for small biology laboratories to become big-data generators. Even if a laboratory is not equipped with such instruments, it can still become a big-data user by gaining access to public repositories containing biological data, such as the US National Center for Biotechnology Information in Bethesda. A large part of the construction in big-data biology is virtual, based on cloud computing, in which data and software are located in massive, off-site centers that can be accessed on demand. Therefore, it is not necessary for users to purchase their own hardware. The cloud computing system allows potential users to create virtual spaces for data, software and results that are freely accessible by everyone, or to keep spaces locked up behind a firewall permitting access to a chosen group of collaborators.
The use of biological databases can be advantageous in several research areas. For example, databases may aid experimental design by allowing the automatic analysis and easy processing of experimental data and making the examination of experimental results simple and quick. Drug discovery is another area that may be simplified by using databases. In this specific area, databases can be scanned in order to find new candidates for drugs by training a classifier on a dataset where functioning and non-functioning drugs have been identified. Moreover, machine learning techniques may be applied to design virtual assays that are able to identify promising new drugs, which can subsequently be analyzed in a laboratory setting. (REF. 4) And most importantly, new scientific experiments can be carried out and new results generated by analyzing existing data sets.
Without the existence of databases, sharing and integration of large quantities of data would be virtually impossible. Although many life scientists have advanced computational skills, a large percentage are not familiar with developing or adapting the relevant software. Nonetheless, the involvement of life scientists in this process is crucial, since they can provide feedback to computer science specialists focusing on different needs and approaches to science. The ability to have access to the actual data sets originally used in a specific study provides researchers with the opportunity to reproduce and expand on such study. This is why it is important for data to become freely available to scientists at any time without restrictions, a notion supported by Open Science and numerous related initiatives. One of these initiatives is known as ELIXIR, a project designed to help scientists across Europe safeguard and share their data and to reinforce current resources, including databases and computing facilities, in individual countries.
Although the creation of biological databases has brought about many benefits, such as the promotion of scientific quality production enabled by networking, they still require improvement in terms of knowledge optimization. It is crucial to manage transdisciplinary knowledge in such a way that will lead to an increase in its quality and quantity. Data heterogeneity is another common issue faced in biological data integration. In the field of biology, several different methods exist for the representation of similar data. This complicates data integration and processing, which, in turn, makes it harder to acquire unified views of such data. An example of this problem is the use of various alternate names when referring to genes, regardless of the existence of full guidelines issued in 1979 proposing the adoption of gene nomenclature standard, leading to difficulties in data sharing. The implementation of standards enables the re-use of data, however, their absence causes significant loss of productivity and contributes to a decrease in data accessible by researchers. Therefore, it is imperative to find a solution to this matter in order to eliminate the challenges faced by scientists when using biological databases to conduct their research.

Final thoughts
Dealing with data implies a drastic discipline to keep access on a long term to the stored information. Technology evolves, which means that the hardware and software used today is not the standard of tomorrow. This means that to be able to read any data written today we will have to execute two different kinds of migrations. A logical migration and a technological migration. Logical migration is related to the kind of format in which the data is stored. Technological migration is related to the kind of hardware used. As an example, if you try to open a Word file written in 1993 with Word version 6 with the latest version Word 2019, it will not work. This example shows a lack of logical compatibility. To avoid this issue and keep an ascended compatibility, the file should have been migrated by the time to the latest version in order to keep it up to date and readable with the latest versions of software.
The same thing applies to hardware, i.e. servers, storage, networks, etc… Another example could be the kind of server and operating system used to run a database. In case you decide to change your hardware and to migrate from, let’s say, Windows to UNIX, a different kind of hardware will be needed to run UNIX and a different version of database to run on UNIX. Windows run on Intel based platforms (and Intel like) and Unix runs on SPARC based platforms, which means that you will have to migrate to a UNIX – SPARC compatible version of the database.
Keeping in mind this constant evolution of hardware, operating systems, software and formats, performing the appropriate logical and technological migrations on time could save you a lot of time and troubles.
Last but not least, it is important to keep backing up your data. Once every three to six months, perform a restore test to see if you are capable of retrieving your backups. This is crucial for two reasons:
- It will keep you up to date on how to restore your data
It is the best testing method to see if your data was properly backed up
Test: LO5 Basic level
References
- Baxevanis AD, Bateman A. 2015. The importance of biological databases in biological discovery. Curr Protoc Bioinformatics., 50(1):1.1.1-1.1.8.
- Benson DA, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW. 2014. GenBank. Nucleic Acids Res., 42:D32–D37.
- Brooksbank C, Bergman MT, Apweiler R, Birney E, Thornton J. 2014. The European Bioinformatics Institute’s data resources 2014. Nucleic Acids Res., 42:D18–D25.
- Caspi R, Billington R, Ferrer L, Foerster H, Fulcher CA, Keseler IM, et al. 2016. The MetaCyc database of metabolic pathways and enzymes and the BioCyc collection of pathway/genome databases. Nucleic Acids Res., 44(D1):D471-80.
- Figueiredo MSN, Pereira AM. 2017. Managing knowledge – the importance of databases in the scientific production. Procedia Manuf., 12:166–73.
- Harris TW, Baran J, Bieri T, Cabunoc A, Chan J, Chen WJ. 2014. WormBase 2014: new views of curated biology. Nucleic Acids Res., 42:D789–D793.
- Howe D, Costanzo M, Fey P, Gojobori T, Hannick L, Hide W, et al. 2008. Big data: The future of biocuration: Big data. Nature., 455(7209):47–50.
- Kanehisa M, Furumichi M, Sato Y, Ishiguro-Watanabe M, Tanabe M. 2021. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res., 49(D1): D545–51.
- Karp PD, Billington R, Caspi R, Fulcher CA, Latendresse M, Kothari A, et al. 2019. The BioCyc collection of microbial genomes and metabolic pathways. Brief Bioinform., 20(4):1085–93.
- Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, Haussler D. 2002. The human genome browser at UCSC. Genome Res., 12(6):996-1006.
- Lapatas V, Stefanidakis M, Jimenez RC, Via A, Schneider MV. Data integration in biological research: an overview. J Biol Res (Thessalon). 2015;22(1):9.
- Marx V. 2013. Biology: The big challenges of big data: Biology. Nature., 498(7453):255–60.
- Nature Structural Biology 10, 980. 2003; doi: 10.1038/nsb1203-980
- Oliveira AL. 2019. Biotechnology, big data and artificial intelligence. Biotechnol J., 14(8):e1800613.
- Razvi SRH, Rampogu S. 2016. Bioinformatics in the present day. MOJ proteom bioinform [Internet]., 3(1):11–2. Available from: http://dx.doi.org/10.15406/mojpb.2016.03.00073
- Toomula N, Kumar A, Kumar D S, Bheemidi VS. 2012. Biological databases- integration of life science data. J Comput Sci Syst Biol., 04(05):087-092. Available from: http://dx.doi.org/10.4172/jcsb.1000081
- Yates AD, Achuthan P, Akanni W, Allen J, Allen J, Alvarez-Jarreta J, et al. 2020. Ensembl 2020. Nucleic Acids Res., 48(D1): D682–8.
- Zou D, Ma L, Yu J, Zhang Z. 2015. Biological databases for human research. Genomics Proteomics Bioinformatics., 13(1):55–63.
- Baxevanis AD, Bateman A. 2015. The importance of biological databases in biological discovery. Curr Protoc Bioinformatics., 50(1):1.1.1-1.1.8.
- Benson DA, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW. 2014. GenBank. Nucleic Acids Res., 42:D32–D37.
- Brooksbank C, Bergman MT, Apweiler R, Birney E, Thornton J. 2014. The European Bioinformatics Institute’s data resources 2014. Nucleic Acids Res., 42:D18–D25.
- Caspi R, Billington R, Ferrer L, Foerster H, Fulcher CA, Keseler IM, et al. 2016. The MetaCyc database of metabolic pathways and enzymes and the BioCyc collection of pathway/genome databases. Nucleic Acids Res., 44(D1):D471-80.
- Figueiredo MSN, Pereira AM. 2017. Managing knowledge – the importance of databases in the scientific production. Procedia Manuf., 12:166–73.
- Harris TW, Baran J, Bieri T, Cabunoc A, Chan J, Chen WJ. 2014. WormBase 2014: new views of curated biology. Nucleic Acids Res., 42:D789–D793.
- Howe D, Costanzo M, Fey P, Gojobori T, Hannick L, Hide W, et al. 2008. Big data: The future of biocuration: Big data. Nature., 455(7209):47–50.
- Kanehisa M, Furumichi M, Sato Y, Ishiguro-Watanabe M, Tanabe M. 2021. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res., 49(D1): D545–51.
- Karp PD, Billington R, Caspi R, Fulcher CA, Latendresse M, Kothari A, et al. 2019. The BioCyc collection of microbial genomes and metabolic pathways. Brief Bioinform., 20(4):1085–93.
- Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, Haussler D. 2002. The human genome browser at UCSC. Genome Res., 12(6):996-1006.
- Lapatas V, Stefanidakis M, Jimenez RC, Via A, Schneider MV. Data integration in biological research: an overview. J Biol Res (Thessalon). 2015;22(1):9.
- Marx V. 2013. Biology: The big challenges of big data: Biology. Nature., 498(7453):255–60.
- Nature Structural Biology 10, 980. 2003; doi: 10.1038/nsb1203-980
- Oliveira AL. 2019. Biotechnology, big data and artificial intelligence. Biotechnol J., 14(8):e1800613.
- Razvi SRH, Rampogu S. 2016. Bioinformatics in the present day. MOJ proteom bioinform [Internet]., 3(1):11–2. Available from: http://dx.doi.org/10.15406/mojpb.2016.03.00073
- Toomula N, Kumar A, Kumar D S, Bheemidi VS. 2012. Biological databases- integration of life science data. J Comput Sci Syst Biol., 04(05):087-092. Available from: http://dx.doi.org/10.4172/jcsb.1000081
- Yates AD, Achuthan P, Akanni W, Allen J, Allen J, Alvarez-Jarreta J, et al. 2020. Ensembl 2020. Nucleic Acids Res., 48(D1): D682–8.
- Zou D, Ma L, Yu J, Zhang Z. 2015. Biological databases for human research. Genomics Proteomics Bioinformatics., 13(1):55–63.
Web sources:
https://en.wikipedia.org/wiki/Airline_reservations_system
https://en.wikipedia.org/wiki/CODASYL
https://en.wikipedia.org/wiki/Database_administrator
https://en.wikipedia.org/wiki/IBM_Information_Management_System
http://www.redbooks.ibm.com/abstracts/sg245352.html
https://en.wikipedia.org/wiki/Navigational_database
https://en.wikipedia.org/wiki/SQL
https://omim.org/
https://www.ascii-code.com
https://www.budapestopenaccessinitiative.org/
https://www.merriam-webster.com/dictionary/data
https://www.snia.org/education/online-dictionary/term/big-data
https://www.snia.org/education/online-dictionary/term/structured-data
https://www.arabidopsis.org/aboutarabidopsis.html


