Introduction to chemistry of proteins

Site: Newgate University Minna - Elearning Platform
Course: General and Medical Biochemistry I
Book: Introduction to chemistry of proteins
Printed by: Guest user
Date: Monday, 24 August 2026, 9:39 PM

1. Proteins Molecules

The word "Protein" was coined by J.J. Berzelius in 1838 and was derived from the Greek word "Proteios" meaning the ‘first rank’ or primary. So, proteins are the major components of any living organism.

Proteins are natural substances with high molecular weights ranging from 5,000 to many millions. Besides Carbon, Hydrogen and Oxygen, they also contain Nitrogen, and sometimes, Sulfur and Phosphorous.

Protein containing foods are essential for living organism, because they are the most important biological molecules in building up and maintenances of the structure of body, giving as much energy as carbohydrates in the course of metabolism in the body. They are also molecular instruments in which genetic information is expressed in living things.

 


1.1. Definition of Proteins

Proteins are macromolecules with a backbone formed by polymerization of amino acids in a polyamide structure. 

Proteins are the most important constituent of cell membranes and cytoplasm. Muscle and blood plasma also contain certain specific proteins.

1.2. Classification of Proteins

There is no universally accepted classification system of proteins, but they may be classified on the basis of the following

•       Composition

•       Solubility

•       Shape

•       Biological function and

•       The three dimensional structure.

 


1.3. Deep dive on the classification of protein

I. Classification base on Composition

A. Simple protein:

Yields only amino acids and no other major organic or inorganic hydrolysis products i.e. most of the elemental compositions. 

B. Conjugated Proteins

These are proteineous molecules that yields amino acids and other organic and inorganic components

E.g., Nucleoprotein (a protein containing Nuclei acids)

 Lipoprotein (a protein containing lipids)

 Phosphoprotein (a protein containing phosphorous)

 Metalloprotein (a protein containing metal ions of Fe2+)

 Glycoprotein (a protein containing carbohydrates).

II. Classification base on Solubility

a) Albumins: These proteins such as egg albumin and serum albumin are readily soluble in water and coagulated by heat.

b) Globulins: These proteins are present in serum, muscle and other tissues and are soluble in dilute salt solution but sparingly in water.

c) Histones: Histones are present in glandular tissues (thymus, pancreas etc.) soluble in water; they combine with nucleic acids in cells and on hydrolysis yield basic amino acids

III. Classification base on shape

A. Fibrous proteins

In this protein, the molecule is constituted by several coiled cross-linked polypeptide chains, they are insoluble in water and highly resistant to enzyme digestion. The ratio of length to breath (axial ratio) is more than 10 in such protein. A few sub groups are listed below.

1. Collagens: the major protein of the connective tissue, insoluble in water, acids or alkalis. But they are convertible to water-soluble gelatin which is easily digestible by enzymes.

2. Elastins: present in tendons, arteries and other elastic tissues, but not convertible to gelatin.

3. Keratins: protein of hair, nails etc.

B. Globular proteins: These are globular or ovoid in shape, soluble in water and constitute the enzymes, oxygen carrying proteins, hormones etc.

Subclasses include: - Albumin, globulins and histones.

IV. Classification base on biological functions:

Proteins are sometimes described as the "workhorses" of the cell because they perform various biological roles.

Enzymes: kinases, transaminases etc.

Storage proteins: myoglobin, ferritin

Regulatory proteins: peptide hormones DNA binding proteins

Structural protein: collagen, proteoglycan

Protective proteins: blood clotting factors, Immunoglobins,

Transport protein: Haemoglobin, plasma lipoproteins

Contractile or motile Proteins: Actin, tubulin


1.4. Classification of proteins base on their level of organization

1. Primary proteins

2. Secondary proteins

3. Tertiary proteins and 

4. Quaternary proteins

a) Primary proteins

The primary structure of a protein is defined by the linear sequences of amino acid residues. Protein contains between 50 and 2000 amino acid residues.

The molecular mass of most proteins is between 5500 and 220,000 Da.

The amino acid composition of a peptide chain has a profound effect on its physical and chemical properties of proteins.

Proteins rich in polar amino acids are more water soluble while proteins rich in aliphatic or aromatic amino groups are relatively insoluble in water and more soluble in cell membranes (can easily cross the cell membrane).

b) Secondary Structure

The secondary structure of a protein refers to the local structure of a polypeptide chain, which is determined by hydrogen bond. 

There are two types of secondary structure, the ∝ - helix and the β- pleated sheet.

The α - helix

The α - helix is a rod like structure with peptide chains tightly coiled and the side chains of amino acid residues extending outward from the axis of spiral. Each amide carbonyl group is hydrogen bonded to the amide hydrogen of a peptide bond that is 4 - residues away along the same chain.

The β- pleated sheet

The β – pleated sheet is an extended structure as opposed the coiled ∝ - helix. It is pleated because the (C-C) bonds are tetrahedral and cannot exist in a planar configuration. If the polypeptide chain runs in the same direction, if forms a parallel β – sheet. It is said to be parallel, and when in opposite direction, antiparallel. 

for example, most immunoglobulins have such β-pleated conformation and some enzymes like Hexokinase contain a mixed α-β conformation.

c) Tertiary Structure

The three dimensional, folded and biologically active conformation of a protein is referred to as tertiary structure. The structure reflects the overall shape of the molecule. The three - dimensional tertiary structure of a protein is stabilized by interactions between side chain functional group, covalent, disulfide bonds, hydrogen bonds, salt bridges, and hydrophobic interactions. 

d)Quaternary Structure

Quaternary structure refers to a complex or an assembly of two or more separated peptide chains that are held together by non-covalent or, in some case, covalent interactions. 

If the subunits are identical, it is a homogeneous quaternary structure; but if there are dissimilarities, it is heterogeneous. For instance, insulin consists of A and B chain which are different. Haemoglobin has 4 chains, two of them are α and two are β. These, the polymers may be dimers, trimers, tetramers and so on.


1.5. Denaturation and factor that affects the denaturation of proteins

Denaturation of Proteins

Proteins have finite lifetimes. They are also subjected to environmental damages like oxidation, proteolysis, denaturation and other irreversible modifications.

Denaturation involves the destruction of the higher-level structural organization (20, 30 and 40) of protein with the retention of the primary structure by the action of denaturing agents.

A denatured protein loses its native physico-chemical and biological properties since the bonds that stabilize the protein are broken down. 

Thus, the polypeptide chain unfolds itself and remain in solution in the unfolded state. The denatured protein may retain its biological activity by refolding (renaturing) when the denaturing agent is removed.

Factors that Affect Denaturation

Denaturing agents

1. physical factors

Temperature, pressure, mechanical shear force, ultrasonic vibration and ionizing radiation causes the protein to lose its biological activity.

2. Chemical factors

Acids and alkalis, organic solvents (actone, ethanol), detergents (cleaning agents), certain amides urea, guandidine hydrochloride, alkaloids, and heavy metal salts (Hg, Cu, Ba, Zn, Cd…) Cause the denaturation.

 Properties of a Denatured Protein

a). An increase in number of reactive and functional group in the composition of the native protein molecule (side chain group of amino acids, COOH, NH2, SH, OH … etc).

B. Reduced solubility and pronounced propensity for precipitation this occurs due to loss of the hydration shell and the unfolding of protein molecules with concomitant exposure of hydrophobic radicals and neutralization of charged polar groups.

C. Configurational alteration of the protein molecule.

D. Loss of biological activity evoked by the disarrangement of the native structural molecular organization.

E. Access of proteolytic enzymes in comparison with the native protein 

Clinical Application of Denaturation

The amounts of proteins found in the urine, serum, CSF are utilized to assess various pathological conditions. The appearance of proteins like Albumin and Globulin in the urine can be detected by precipitating them using ammonium sulphate. 

This could be used to asses the degree of kidney impairment and glomerular permeability.

In some disease, abnormal proteins may be present in plasma and be filtered at the glomeruli. The most important member is Bence-jons’ protein which is most often associated with multiple myeloma. 

So, recognition of such protein in the urine may be useful in the diagnosis of the disease.


1.6. Digestion and Absorption of Proteins

Digestion and Absorption of Proteins

Proteins are large polypeptide molecules coiled by weaker bonds in their tertiary structure, the digestion of proteins involves the gradual breakdown of this polypeptide by enzymatic hydrolysis into amino acid molecules which are absorbed in the blood stream. 

The protein load received by the gut is derived from two sources 70-100g dietary protein which is required daily and 35 - 200g endogenous protein (secreted enzymes and proteins in the gut or from intestinal epithelia cell turnover).

Only 1-2g of nitrogen equivalent to 6-12g of proteins are lost in the faeces on a daily basis. 

The process of protein digestion can be divided, depending on the sources of peptidases.

A. Gastric Digestion

Entry of a protein in to stomach stimulates the gastric mucosa to secrete a hormone gastrin which in turn stimulates the secretion of HCl by the parietal cells of the gastric glands and pepsinogen by the chief cells.

The HCL thus produced lower the pH of stomach to (pH 1.5 – 2.5) and acts as an antiseptic and kills most of the bacteria and other foreign cells ingested along with.

The acid denatures the protein and the whole protein susceptible to hydrolysis by the action other proteolytic enzymes.

Proteases are endopeptidases which attack the internal bonds and liberate large fragments of peptides.

Then pepsinogen having MW 40,000 an inactive precursor or zymogen is converted in to active pepsin in the stomach itself. In this process 44 amino acids gets removed from the amino terminal end and the portion of the molecule that remain intact is enzymatically active pepsin (MW. 33,000).

This active pepsin cleaves the ingested protein at their amino terminus of aromatic amino acids (Phe, Tyr, and Trp.). The major products of pepsin action are large peptide fragments and some free amino acids.

B. Pancreatic Digestion

Pancreatic zymogens proceed digestion as the acidic stomach contents pass in to the small intestine, A low pH triggers the secretion of a hormone Secretin in the blood. Secretin stimulates the pancreas to secrete HCO3- (bicarbonate), which in the small intestine neutralizes the gastric HCL and abruptly change the pH to 7.0.

The entry of large peptide fragments and some free amino acids in the upper part of the small intestine (Duodenum), excites the release of a hormone cholecystokinin (CCK).

C. Intestinal Digestion

Since pancreatic juice does not contain appreciable aminopeptidase activity final digestion of di and Oligopeptides depends on the small intestinal enzymes.

The lumenal surface of epithelial cells is rich in endopeptidase, and dipeptidase aminopeptidase activity.

The end products of the cell surface digestion are free amino acids and di and tripeptides. These are passed into the interior of the epithelial cell where other specific peptidases convert almost all of them to a single amino acid that are transported to the blood stream by the opposite side of the cell membrane and carried to liver (primarily) and other tissues for oxidative degradation. This process completes the absorption of 99% of digested proteins.

 


1.7. Amino acids

Introduction

There are approximately 300 amino acids present in various animals, plants, and microbial systems, but only 20 amino acids are coded by DNA to appear in proteins.

Cells produce proteins with strikingly different properties and activities by joining the same 20 amino acids in many different combinations and sequences. This indicates that the properties of proteins are determined by the physical and chemical properties of their monomeric units, the amino acids.

Definition

Amino acids are the basic structural units of proteins consisting of an amino group, (-NH2) a carboxyl (-COOH) group a hydrogen (H) atom and a (variable) distinctive (R) group.

All of the substituents in amino acid are attached (bonded) to a central α carbon atom. This carbon atom is called α because it is bonded to the carboxyl (acidic) group.


General structure of amino acids


1.8. Stereochemistry (Optical activity)

Stereochemistry (Optical activity)

Stereochemistry mainly emphasizes the configuration of amino acids at the α carbon atom, having either D or L- isomers.

      COOH                                         COOH

       |                                                     |

 H - C – NH2                             H2N – C – H

       |                                                     |

      R                                                    R

D (+) amino acid                      L (-) amino acid

D and L- forms of Amino acids

Out of the 20 amino acids, proline is not an α amino acid rather an α - imino acid. Except for glycine, all amino acids contain at least one asymmetric carbon atom (the α - carbon atom).


1.9. Classification of Amino acids

Classification of Amino Acid

L-Amino acids are the building blocks of proteins. They are frequently grouped according to the chemical nature of their side chains. Links to individual amino acids are given below:

I. Structural Classification

This classification is based on the side chain radicals (R-groups). Each amino acid is designated by three letter abbreviation e.g. Aspartate as Asp and by one letter symbol D.

II. Electrochemical classification

Amino acids could also be classified based on their acid – base properties

 

 Acidic amino acids (Negatively charged at pH = 6.0)

 Example:

  • Aspartic acid               - CH2 – COO-
  • Glutamic acid              - CH2 – CH2 – COO

Basic amino acids (positively charged at PH = 6.0)

Example:

  • Lysine                          - CH2 – CH2 - CH2 - CH2 –NH3+
  • Neutral amino acid

 Example:

♦ Serine - CH2- OH

♦ Threonine - CH2- OH

                          ↓

                         CH3

 ♦ Asparagine - CH2- CO-NH2

 ♦ Glutamine - CH2- CH2 - CO-NH2

 

Biological/Physiological Classification

This classification is based on the functional property of amino acids for the organism.

1. Essential Amino Acids

Amino acids which are not synthesized in the body and must be provided in the diet to meet an animal’s metabolic needs are called essential amino acids. 

About ten of the amino acids are grouped under this category indicating that mammals require about half of the amino acids in their diet for growth and maintenance of normal nitrogen balance. 

2. Non- Essential Amino Acids

These amino acids are need not be provided through diet, because they can be biosynthesized in adequate amounts within the organism.

Essential Amino Acids in Mammals

Arginine, Histidine, Isoleucine, Leucine, Lysine, Methionine, Phenylalanine, Threonine, Tryptophan, Valine

 Non-Essential Amino Acids in Mammals

Alanine, Asparagine, Aspartic Acid, Cysteine, Glutamic Acid, Glutamine, Glycine, Proline, Serine, Tyrosine

3. Semi-essential amino acids

Two amino acids are grouped under semi-essential amino acids since they can be synthesized within the organism but their synthesis is not in sufficient amounts. In that they should also be provided in the diet.

The set of essential amino acids required for each species of an organism can be an indicative of the organism propensity to minimal energetic losses on the synthesis of amino acids. Semi essential amino acids include Arginine and Histidine

IV. Classification Based on the Fate of Each Amino acid in Mammals.

Amino acids can be classified here as Glucogenic (potentially be converted to glucose), ketogenic (potentially be converted to ketone bodies) and both glucogenic and ketogenic.

I. Glucogenic Amino Acids

Those amino acids in which their carbon skeleton gets degraded to pyruvate, α ketoglutarate, succinyl CoA, fumarate and oxaloacetate and then converted to Glucose and Glycogen, are called as Glucogenic amino acids.

They include: -

Alanine, cysteine, glycine, Arginine, glutamine, Isoleucine, tyrosine.

II. Ketogenic Amino Acids

Those amino acids in which their carbon skeleton is degraded to Acetoacetyl CoA, or acetyl CoA. then converted to acetone and β-hydroxy butyrate which are the main ketone bodies are called ketogenic amino acids.

They include: -

Phenylalanine, tyrosine, tryptophan, isoleucine, leucine, and lysine. 

These amino acids have ability to form ketone bodies which is particularly evident in untreated diabetes mellitus in which large amounts of ketone bodies are produced by the liver (i.e. not only from fatty acids but also from ketogenic amino acids)

Degradation of Leucine which is an exclusively ketogenic amino acid makes a substantial contribution to ketone bodies during starvation.

III. Ketogenic and Glucogenic Amino Acids

The division between ketogenic and glucogenic amino acids is not sharp for amino acids (Tryptophan, phenylalanine, tyrosine and Isoleucine are both ketogenic and glucogenic).

Some of the amino acids that can be converted in to pyruvate, particularly (Alanine, Cysteine and serine, can also potentially form acetoacetate via acetyl CoA especially in severe starvation and untreated diabetes mellitus.


Answer this**
A child with tall stature, loose joints, and detached retinas is found to have a mutation in collagen. Which of the following amino acids is the recurring amino acid most likely to be altered in mutations that distort collagen molecules?