Background: Cases A, B and C deal with PKU, a genetic metabolic disease caused by a mutation in the phenylalanine hydroxylase enzyme (PAH). In the most common form, a C to T point mutation (SNP) causes an arginine to be replaced by tryptophan at amino acid position 408, resulting in an inactive enzyme and incomplete metabolism of phenylalanine-containing compounds such as proteins. The NCBI designator for this SNP is rs5030858 (used in Step 1 below). There are a variety of ways to obtain information about this SNP using NCBI tools, three of which will be demonstrated below (Genome Data Viewer, dbSNP, and Structure).
Goal of Steps 1-7: Use the Genome Data Viewer and SNP database to learn more about SNPs associated with PKU.
Goal of Steps 8-15: Use the NCBI Structure tool to see the replacement of arginine by tryptophan in 3 dimensions.
1. Open the Genome Data Viewer and enter rs5030858 into the search box*, as shown below, and click the magnifying glass icon so search for this SNP.
*Note: If you get an error message when attempting to search for rs5030858 from the GDV home page, type in the name of the gene instead (PAH), which will take you to the next screen. Then type rs5030858 into the search field of that screen (see upper left of image in Step 2 below).
Note: Use the keystroke combinations control-C (or command-C) and control-V (or command -V) to copy and paste, and use the backspace key (Windows keyboard) or delete key (Mac keyboard) f you need to delete text from a search field.
2. The location of the SNP is shown on the gray double-stranded DNA sequence, which is also represented by the top green line. Click on the top green line and the amino acid sequences of the protein will appear, represented by the red line. The direction of arrows point from right to left, indicating the 5′ to 3′ direction of the DNA molecule, so the nucleotides in the red box should be read from right to left as “CGG”, a codon for arginine (R). Hover your mouse over the top green line, in the location of the red box on that line, and a pop-up box will appear giving the name of the gene (PAH) and its associated protein (phenylalanine hydroxylase). There are 13 small circles in the blue bar, representing the 13 exons of the PAH gene. The 12th circle from the left is open, indicating that this DNA sequence is part of exon 12.
3. Move your mouse cursor down and hover over the red box labelled rs5030858, as shown below, and click on the rs5030858 link next to SNP summary in the pop-up box.
4. The SNP database entry will appear, with “Alleles G>A” indicated as shown below. Click on the Variant Details tab and scroll down to PAH transcript variant 1 (red box at bottom of image below). [CGG] > W [TGG] is listed under Amino acid [codon] for this variant. If this is a C to T mutation, why does it say “Alleles G>A”?
For the answer, look at the “wild type” (non-mutated) DNA in the image above, and note that it has a G-C base pair at the SNP location. If the G is changed to an A because of the mutation, then the A would pair with T so there would now be an A-T base pair in the mutated DNA. The PAH gene is coded on the reverse (bottom) strand, so (reading from right to left on the bottom strand) the DNA codon would change from CGG (arginine, symbol R) to TGG (tryptophan, symbol W. That is why this this mutation is referred to in the literature as the c.1222C>T mutation. The number 1222 refers to the CDS (coding position) of the base pair running from the start codon to the stop codon. This corresponds to amino acid position 408 (122/3 = 408). Because of this, it can also be called the p.Arg408Trp or R408W mutation.
5. Click on the Clinical Significance tab and scroll down to another view of the gene location. Mouse over the box labelled 577 (which lines up with rs5030858) and note that the rs5030858 SNP also has a ClinVar Variation ID: 577, indicating that this SNP has been evaluated as a potential pathogen. It is the most common cause of PKU, but there are hundreds of others in PAH gene. Consequently, in the United States PKU patients are likely to be compound heterozygotes. This means that the patient has two recessive alleles for the same gene, but the alleles are different (for example, they are found at different locations).
6. To see another example, click on the box labelled 612, just to the left of 577 box, then mouse over the 612 box and click the SNP ID rs5030859.
7. This SNP mutation is also associated with PKU, as you can see by clicking the Clinical Significance tab. Click the Variant Details tab to see the three variations that can occur in this location.
Question: How many other SNPs can you find associated with PKU or other clinical conditions?
Goal of Steps 8-15: Use the NCBI Structure tool to see the replacement of arginine by tryptophan in 3 dimensions.
8. Open NCBI Structure, enter human phenylalanine hydroxylase into the search box, and click Search. If you copy and paste, use control-V or command -V to paste.
9. Click the view in iCn3D link as shown in the red box below.
10. The phenylalanine hydroxylase (PAH) protein structure can be rotated in three dimensions by dragging with your left mouse button, and the entire structure can be moved using by right-clicking and dragging. Use the wheel of your mouse to make the structure larger or smaller. Click the ClinVar checkbox, then click the Details tab as shown below.
Note: On a tablet or phone, use one finger to rotate the structure, and two fingers to move the entire structure. If the structure disappears when you attempt to manipulate it, hit the back arrow of your browser to go back to Step 9, and click on View in iCn3D again. You may have to experiment with finger movements, especially on a phone.
11. Drag the bottom scrollbar to the right to see position 408 on the protein, and hover over the W using your mouse cursor until the gray box pops up. Click on the 3D with scap button.
Note: On a phone or tablet, carefully drag the protein sequence with your figure to see position 408, on the far-right end of the protein. Delicately touch the W with your finger until the pop-up box appears.
12. Drag the left edge of the Sequences and Annotations window to move it to the right (or drag the entire window by its title bar) to make the protein structure easier to see. Press the A key on your keyboard (if using a computer) or use the View menu (if using a tablet or phone). Either method will allow you to alternate between seeing arginine or tryptophan at position 408, in three dimensions.
Note: On a phone, the View menu can be accessed via a menu button with three lines at upper left, not shown here.
13. Pressing the A key or using the View menu will allow you to toggle between arginine on the wild type PAH (shown directly below) and tryptophan on the mutant PAH (second image below). Rotate the structure by dragging with your mouse, and enlarge using your mouse wheel (or the pinch method on a phone or tablet). The label designations shown below refer to arginine or tryptophan at position 408 of the protein. Click here to see a video showing how the structure can be rotated to identify the adjacent amino acids (proline in both cases).
14. Drag the left edge of the Sequences and Annotations window back to the left, hover over the W again at position 408, but this time click the Interactions button in the pop-up box.
15. Use your mouse and mouse wheel (or the pinching method) to rotate and enlarge the protein to see the interactions between the amino acids and neighboring amino acids. Press the A key or using the View menu will allow you to toggle between arginine (shown directly below) and tryptophan (second image below).
Note: If you are looking at interactions on the wild type protein, pressing the A key may not show them on the mutant. You may need to hit your back browser arrow, go back to Step 9 -10, then skip to Step 14. You should then be able to toggle back and forth and see the interactions with the arginine and tryptophan in different colors than the surrounding amino acids.
Questions:
1. How might changes in these interactions interfere with protein (enzyme) function?
2. What is the function of phenylalanine hydroxylase, and why is it so important that the body be able to metabolize phenylalanine-containing compounds?
















