Biological chemistry in water
Structures and applications
Proteins, carbohydrates, lipids, and nucleic acids are built from common organic functional groups. What is new is the solvent: water, which affects the acid-base chemistry of these structures and, through the hydrophobic effect, affects how these molecules fold and assemble.
Water as the reaction medium
Hydrogen bonding, the hydrophobic effect, and physiological pH
Every reaction in a cell happens in water, and water is a participating solvent. Each molecule can donate two hydrogen bonds and accept two, so liquid water is an extensively hydrogen-bonded network and very polar. Any solute placed in that network has to interact with water, and polarity and hydrogen bonding potential decides what dissolves, what folds, and what assembles.
A polar or charged group is hydrated by an ordered shell of water molecules, which is why ions, sugars, and amides are water soluble. A non-polar hydrocarbon group cannot donate or accept a hydrogen bond, so the water around it must reorganize into a more ordered cage; that loss of entropy is unfavourable, and the system responds by clustering the nonpolar groups together to minimize the surface exposed. This hydrophobic effect is entropic in origin, not necessarily a special attraction between hydrocarbons, and it is the main force behind protein folding and membrane formation.
Water is also a weak acid and a weak base, so it dictates what acid–base chemistry is possible. At physiological pH near 7.4, a carboxylic acid with a pKa near 4 is fully deprotonated and an aliphatic amine with a pKa near 10 is fully protonated. Such biological molecules therefore mostly exist as ions, and comparing a group’s pKa with the local pH is the first thing to do before drawing any structure or mechanism.
Amino acids and proteins
Zwitterions, the peptide bond, and why a chain folds
An α-amino acid carries an amino group and a carboxylic acid on the same carbon, and in water that pairing means the acid has donated its proton to the amine. The neutral species is therefore a zwitterion with a positive ammonium and a negative carboxylate, which is why amino acids themselves are high-melting crystalline solids. The pH at which the net charge is zero is the isoelectric point, and it depends on the side chain attached.
The nineteen chiral amino acids all occur in Nature as the L enantiomer, and their attached side chains are what distinguish them: nonpolar chains that aggregate in the interior of a folded protein, polar and hydrogen-bonding chains that sit on the surface interacting with water, and ionizable chains that carry charge. Histidine matters especially because its imidazole pKa is near 6, close enough to physiological pH that it can act as either acid or base, which is why it appears in so many enzyme active sites.
Condensation of the carboxylate of one residue with the amine of the next gives the amide link in a peptide bond. Amide resonance gives that bond substantial double-bond character, so it is planar, restricted in rotation, and slow to hydrolyze; the polypeptide backbone is consequently a fairly rigid ribbon whose only real freedom is rotation about the two single bonds flanking each α-carbon. Backbone hydrogen bonding then builds the α-helix and β-sheet, and side-chain burial by the hydrophobic effect drives the whole chain into one tertiary structure.
In a cell, amino acid condensation is completely selective, since the chemistry is catalyzed by enzymes, which select for exactly one amine nucleophile and one carboxylate electrophile. In a lab flask it is not selective: two amino acids each carrying both groups can couple four different ways. Synthetic peptide chemistry therefore relies on protecting group chemistry, blocking the amine of one partner as a carbamate such as Boc or FMoc and the acid of the other as an ester, so that only the intended pair can react. Removing each group under selective conditions that leave the other untouched allows for the synthesis of a defined peptide sequence.
Carbohydrates
Cyclic hemiacetals, anomers, and the glycosidic bond
A monosaccharide is a polyhydroxy aldehyde or ketone, and in water very little exists as the open-chain form. The hydroxyl on C5 of a hexose adds in an intramolecular sense to the C-1 carbonyl to give a six-membered cyclic hemiacetal, a pyranose, and because that addition creates a new stereocentre at the former carbonyl carbon there are two outcomes. That carbon is the anomeric centre, and the α and β anomers interconvert freely through the open-chain form, a process known as mutarotation.
Which anomer predominates is a conformational and stereoelectronic question. In glucose the β-anomer places every substituent, including the anomeric hydroxyl, in an equatorial position on the ring chair, contributing to why glucose is the most abundant hexose in Biology. The anomeric hydroxyl is also the most reactive one, since it forms an acetal on reaction with another alcohol; that acetal is known as the glycosidic bond.
Linking monosaccharides through glycosidic bonds gives structures ranging from sucrose to cellulose, and the stereochemistry of the link determines the biological function. Starch uses α-1,4 links and coils into a helix that mammalian enzymes hydrolyze easily; cellulose uses β-1,4 links, which give an extended ribbon held together by hydrogen bonds into fibres that humans cannot digest. Two polymers made from the same monomer behave completely differently because of one stereocentre.
A sugar whose ring can open to expose a free aldehyde or ketone is a reducing sugar, as it is reduced by sodium borohydride and oxidized by a wide range of oxidants, while a glycoside, whose anomeric position is an acetal, cannot open and is not reduced or oxidized. The orientation of the hydroxyl at the last stereocentre of the open-chain form defines the D or L series, with most biological sugars being D.
Lipids and membranes
Amphiphiles, self-assembly, and the bilayer
Lipids are grouped by solubility rather than by a common functional group, and the ones that matter structurally are the amphiphiles: a molecule with a polar head and one or more long hydrocarbon tail. A triglyceride, three fatty acids linked to glycerol through ester bonds, is a storage molecule and is not amphiphilic; hydrolysis of it with aqueous base gives the carboxylate salts we call soap, and those are very useful amphiphiles.
Put an amphiphile in water and the hydrophobic effect dictates how it behaves. At low concentration the tails aggregate inward and the heads face outward towards the solvent, giving a spherical micelle. A phospholipid with two tails packs better as a flat sheet, and two such sheets tail-to-tail give the bilayer that encloses every cell. Nothing directs this assembly but the the interactions with water, and the structure is held together by many weak interactions rather than covalent bonds, which is why membranes are fluid and self-healing.
The degree of unsaturation in the tails sets that fluidity. A saturated chain packs tightly and raises the melting point, which is why animal fats are solid; a cis double bond puts a permanent bend in the chain, disrupts packing, and keeps vegetable oils and membranes liquid at room temperature. Cholesterol, a rigid fused-ring alcohol, inserts between the tails and buffers the membrane against changes in either direction.
Nonpolar molecules serve many useful purposes. Arachidonic acid, a C20 acid with four cis double bonds, is oxidized and cyclized in the body to the prostaglandins, potent signalling molecules whose synthesis is the target for aspirin and the other non-steroidal anti-inflammatories. Conversely, the triglycerides are pure storage, and both they and the fatty acids are assembled iteratively, two carbons at a time from acetyl-CoA.
The terpenes come from the same iterative building with small repeating units, in their case five-carbon isoprene groups. Geraniol is two of them attached; joining and cyclizing more gives, for example, camphor, longifolene, taxadiene, and, after further elaboration, cholesterol and the steroid hormones. Recognizing the isoprene pattern in a complicated skeleton is often the fastest way to work out how it was made in Nature.
Nucleic acids
Nucleotides, the phosphodiester backbone, and base pairing
A nucleotide has three parts: a heterocyclic base, a ribose or deoxyribose sugar attached at the anomeric carbon through an N-glycosidic bond, and one to three phosphate groups attached at C5′. The sugar is a five-membered furanose, ribose in RNA and 2′-deoxyribose in DNA, and that single hydroxyl is the difference between the two polymers. The phosphates are fully ionized at physiological pH, so nucleic acids are polyanions, and that charge is why they need counterions around them.
Linking the 3′-hydroxyl of one nucleotide to the 5′-phosphate of the next gives the phosphodiester backbone and defines a direction along the chain. Both polymers use the purines adenine and guanine, while the pyrimidines differ; thymine and cytosine in DNA, uracil and cytosine in RNA. The bases then pair with hydrogen bonds in one specific way; adenine with thymine, or with uracil in RNA, through two hydrogen bonds, and guanine with cytosine through three. The pairing is dictated by which donors and acceptors line up across the helix, and it is what makes the sequence of one strand a template for the other.
Two antiparallel strands wind into a double helix with the ionic backbone outside in contact with water and the flat aromatic bases stacked inside away from it. Hydrogen bonding gives the specificity, but the stacking of the bases and the exclusion of water from that interior give most of the stability. Heating separates the strands, and because G–C pairs contribute an extra hydrogen bond, a GC-rich sequence melts at a higher temperature.
Enzymes and biological catalysis
Binding, the transition state, and the same mechanisms you already know
An enzyme is a catalyst, and it lowers the activation barrier without changing the position of the equilibrium. It does so by binding the transition state more tightly than it binds the substrate. The active site is a pocket shaped by the folded chain that holds the substrate in one orientation, excludes water where water would interfere, and places acidic and basic side chains exactly where a mechanism needs them. That precise arrangement is what causes the enormous rate accelerations.
The chemistry itself involves typical organic processes. Serine proteases use an alcohol made nucleophilic by a histidine and an aspartate to attack an amide carbonyl through a tetrahedral intermediate; that is nucleophilic acyl substitution. Aldolases make and break carbon–carbon bonds through an enamine or enolate nucleophile; that is enolate chemistry. Dehydrogenases move a hydride to NAD+; that is a redox step involving nucleophilic addition. Every one of these is a reaction from an earlier chapter, catalyzed by the enzyme template.
Reference
Reaction summary
Each transformation and structural relationship in this chapter.
Self-check
Six questions before you move on
Try to work out an answer on paper, then reveal to check. If your reasoning is right but the answer is wrong, keep going.
Nonpolar groups cluster together in water. Is that because hydrocarbons attract one another strongly?
No. The London forces between hydrocarbon chains are weak. The driving force sits with the water; a nonpolar surface forces the surrounding water into a more ordered, hydrogen-bonded cage, which costs entropy. Clustering the nonpolar groups reduces the total surface that has to be caged, so the entropy of the system increases. The hydrophobic effect is a property of the solvent, not an attraction between the solutes.
Glycine has a molecular weight of 75, yet it melts above 230 °C rather than boiling like a small organic molecule. Why?
Because it is not neutral. The carboxylic acid protonates the amine in the same molecule, so glycine exists as a zwitterion with a full positive and a full negative charge. The solid is held together by ionic attractions rather than by dipole forces, so it behaves like a salt: high melting, water soluble, and insoluble in nonpolar solvents.
Starch and cellulose are both polymers of glucose joined 1,4. Why can you digest one and not the other?
The anomeric stereochemistry. Starch uses α-1,4 links, which bend the chain into a helix that our amylase enzymes are shaped to bind and hydrolyze. Cellulose uses β-1,4 links, which give a flat extended ribbon that hydrogen bonds to its neighbours into tough crystalline fibres, and we have no enzyme with an active site that fits it. One stereocentre separates digestible food from fibre.
A single-tailed fatty acid salt forms micelles, but a two-tailed phospholipid forms bilayers. What decides which?
Shape, and how the molecules pack. A single tail gives a cone-shaped molecule with a wide head and narrow tail, and cones pack into a sphere. Two tails make the cross-section of the tail region roughly the same as the head group, giving a cylinder, and cylinders pack into a flat sheet. Two sheets placed tail-to-tail bury all the hydrocarbon and expose only heads to water.
A DNA sequence rich in G and C melts at a higher temperature than one rich in A and T. Explain.
G pairs with C through three hydrogen bonds, while A pairs with T through only two, so each GC pair takes more energy to break. GC pairs also stack somewhat more favourably. More heat is therefore needed to separate the strands, which is why GC content is used to predict the melting temperature of a duplex.
Where in an enzyme mechanism does new chemistry appear, compared with the reactions you learned earlier?
Nowhere. The steps are the same ones: nucleophilic acyl substitution in a protease, enolate chemistry in an aldolase, hydride transfer in a dehydrogenase. What the enzyme adds is not a new reaction but templating: it holds the substrate in one orientation, supplies acidic and/or basic groups exactly where they are needed, excludes water, and stabilizes the transition state. That is why the rate rises enormously while the equilibrium does not shift.