Learn C Programming

Lesson 4 of 8 ¡ Meet C: Your First Programs, Symbol by Symbol

Module 1 ¡ Meet C: Your First Programs, Symbol by Symbol

The Complete ASCII Table, All 128 Rows

FreeReading

In this lesson

  • Explain why a char in C is really a small integer.
  • Recall the four anchors: '0' is 48, 'A' is 65, 'a' is 97, a space is 32.
  • Use the three arithmetic tricks, and read any row of the complete table below.

A computer stores numbers. Only numbers. So how does it store the letter A?

It agrees on a code. In 1963 a standard wrote the agreement down: A is 65, B is 66, and so on to 127. Its name is the American Standard Code for Information Interchange, ASCII for short. Every machine that follows it reads the same byte the same way.

This lesson is a reference page with the whole agreement printed in it, and five short programs that use it.

One byte, two readings

In C, a char is not really a letter. It is a small integer that you have agreed to show on the screen as a letter.

That is why the same variable can be printed two ways. %c says "show me the character this number stands for". %d says "show me the number".

#include <stdio.h>

int main(void)
{
    char letter = 'A';

    printf("As a character: %c\n", letter);
    printf("As a number:    %d\n", letter);
    return 0;
}
As a character: A
As a number:    65

Nothing was converted between those two lines. The byte in memory never changed. Only the lens you looked at it through changed.

The four anchors and the three tricks

You do not need to learn 128 numbers. You need four, and everything else follows by arithmetic.

CharacterCodeThe whole group
'0'48The digits '0' to '9' are 48 to 57
'A'65The capitals 'A' to 'Z' are 65 to 90
'a'97The small letters 'a' to 'z' are 97 to 122
a space32The first printable character of all

From those four, three tricks cover most of the real work you will do with characters.

  1. Change case by 32. A small letter is exactly 32 above its capital, so 'a' - 'A' is 32.
  2. Digit character to digit value: subtract '0'. '7' - '0' is 7, because 55 minus 48 is 7.
  3. Letter to alphabet position: subtract 'A' or 'a'. 'D' - 'A' is 3, so D is the fourth letter.

Notice what all three have in common. You never write 48, 65 or 97 in your code. You write '0', 'A' or 'a', and the compiler fills the number in, which keeps the line readable.

So if you can subtract, you can do case conversion, digit parsing and alphabet positions without a single library function.

The control characters: 0 to 31, and 127

The first 32 codes are not printable. They are control characters: commands left over from teleprinters, machines that typed text down a wire. Most are museum pieces. Five still matter, and they are in bold.

DecHexNameFull nameWhat it does
000NULNullEnds a C string. The one you will meet in Module 10
101SOHStart of headingA teleprinter header marker, unused today
202STXStart of textMarked where the message began
303ETXEnd of textWhat Ctrl and C send to a terminal
404EOTEnd of transmissionWhat Ctrl and D send, read as end of input
505ENQEnquiryAsked the other machine to answer
606ACKAcknowledgeThe other machine's yes
707BELBellRings the terminal bell. Written \a
808BSBackspaceMoves back one character. Written \b
909HTHorizontal tabJumps to the next tab stop. Written \t
100ALFLine feedThe newline on Linux and macOS. Written \n
110BVTVertical tabMoved down a fixed distance on paper
120CFFForm feedEjects a page on a printer. Written \f
130DCRCarriage returnBack to the start of the line. Windows writes CR then LF
140ESOShift outSwitched to an alternative character set
150FSIShift inSwitched back again
1610DLEData link escapeTold the link that what follows is a command
1711DC1Device control 1Also called XON: resume sending
1812DC2Device control 2A spare device signal
1913DC3Device control 3Also called XOFF: stop sending
2014DC4Device control 4A spare device signal
2115NAKNegative acknowledgeThe other machine's no
2216SYNSynchronous idleFilled a quiet line to keep the timing
2317ETBEnd of transmission blockEnded one block of a longer message
2418CANCancelSaid the data before it was wrong
2519EMEnd of mediumThe tape or the card ran out
261ASUBSubstituteCtrl and Z on Windows, read as end of file
271BESCEscapeStarts a terminal colour code. Written \033
281CFSFile separatorThe coarsest of the four separators
291DGSGroup separatorSeparated groups inside a file
301ERSRecord separatorSeparated records inside a group
311FUSUnit separatorSeparated fields inside a record
1277FDELDeleteAll seven holes punched on paper tape, which meant ignore this

The five in bold are the ones you will actually type. NUL ends a string, HT is \t and LF is \n. CR is the other half of a Windows line ending, and ESC starts the codes that colour a terminal.

So a text file with no control characters in it at all would be one endless line.

The printable half: 32 to 126, complete

Everything from 32 to 126 has a shape you can see. Read the table in four blocks, left to right. The first cell of the first block is the space character, which is why it looks empty.

DecHexCharDecHexCharDecHexCharDecHexChar
3220space563888050P10468h
3321!573998151Q10569i
3422"583A:8252R1066Aj
3523#593B;8353S1076Bk
3624$603C<8454T1086Cl
3725%613D=8555U1096Dm
3826&623E>8656V1106En
3927'633F?8757W1116Fo
4028(6440@8858X11270p
4129)6541A8959Y11371q
422A*6642B905AZ11472r
432B+6743C915B[11573s
442C,6844D925C\11674t
452D-6945E935D]11775u
462E.7046F945E^11876v
472F/7147G955F_11977w
483007248H9660`12078x
493117349I9761a12179y
50322744AJ9862b1227Az
51333754BK9963c1237B{
52344764CL10064d1247C|
53355774DM10165e1257D}
54366784EN10266f1267E~
55377794FO10367g

Three patterns are worth seeing in the table itself. The digits sit together, the capitals sit together, and the small letters sit together, each group unbroken.

Look at 65 and 97. The gap between a capital and its small letter is 32 on every row, which is trick one, visible.

Look at 48 and 57. The digits are in order with no gaps, which is what makes trick two work for every digit and not just for '7'.

Beyond ASCII, and what that means for Bangla

ASCII only defines 0 to 127, because it was designed for seven bits. A byte holds eight, so codes 128 to 255 were left free and every vendor filled them differently. That mess is called "extended ASCII" and it is why the same file looked like nonsense on a different machine.

Unicode fixed it by giving every character in every script its own number, and UTF-8 is the way those numbers are stored as bytes. ASCII is the first 128 of them, unchanged, which is why old files still work.

Here is the part that matters for you. In UTF-8, one Bangla letter takes three bytes, not one. The word "āĻĸāĻžāĻ•āĻž" is four letters and twelve bytes, and strlen answers 12.

So a char holds one byte, not one letter. The three tricks work on ASCII letters and digits. They do not work on Bangla text, because there is no single number to add 32 to. Real Bangla text handling is a Module 10 topic and needs a library.

Example 1: the same byte, four ways

The smallest program that shows the whole idea. One variable, four printed lines, no conversion anywhere.

#include <stdio.h>

int main(void)
{
    char letter = 'A';

    printf("As a character: %c\n", letter);
    printf("As a number:    %d\n", letter);
    printf("Add 1:          %c\n", letter + 1);
    printf("Add 32:         %c\n", letter + 32);
    return 0;
}
As a character: A
As a number:    65
Add 1:          B
Add 32:         a

Adding 1 moved along the alphabet. Adding 32 jumped from the capitals to the small letters. Both are ordinary arithmetic on the number 65.

Run in Compiler
Example 2: change case without a library

Trick one, in both directions. Add 32 to go down to small letters, subtract 32 to go up to capitals.

#include <stdio.h>

int main(void)
{
    char upper = 'M';
    char lower = 'q';

    /* A small letter is exactly 32 above its capital. */
    printf("%c becomes %c\n", upper, upper + 32);
    printf("%c becomes %c\n", lower, lower - 32);
    return 0;
}
M becomes m
q becomes Q

This only works if the character really is a letter. Try it on '5' and you get code 21, an invisible control character, with no warning at all.

Run in Compiler
Example 3: a digit character is not its value

Trick two, and the mistake it exists to prevent. The character '7' is the number 55, and the number 7 is something you have to work out.

#include <stdio.h>

int main(void)
{
    char digit = '7';

    int wrong = digit;           /* 55, the ASCII code */
    int right = digit - '0';     /* 7, the value it stands for */

    printf("Wrong: %d\n", wrong);
    printf("Right: %d\n", right);
    printf("Squared: %d\n", right * right);
    return 0;
}
Wrong: 55
Right: 7
Squared: 49

That - '0' appears in almost every program that reads digits out of text, which is most programs that read anything.

Run in Compiler
Example 4: print the printable half yourself

This one uses a loop, which is Module 6. Read it, run it, and do not worry about writing your own yet. The point is that the table above is not a fact to trust, it is something your machine can print for you.

#include <stdio.h>

int main(void)
{
    for (int code = 32; code <= 126; code++) {
        printf("%3d %c", code, code);

        /* Start a new line after every eighth entry. */
        if ((code - 31) % 8 == 0) {
            printf("\n");
        } else {
            printf("   ");
        }
    }
    printf("\n");
    return 0;
}
 32      33 !    34 "    35 #    36 $    37 %    38 &    39 '
 40 (    41 )    42 *    43 +    44 ,    45 -    46 .    47 /
 48 0    49 1    50 2    51 3    52 4    53 5    54 6    55 7
 56 8    57 9    58 :    59 ;    60 <    61 =    62 >    63 ?
 64 @    65 A    66 B    67 C    68 D    69 E    70 F    71 G
 72 H    73 I    74 J    75 K    76 L    77 M    78 N    79 O
 80 P    81 Q    82 R    83 S    84 T    85 U    86 V    87 W
 88 X    89 Y    90 Z    91 [    92 \    93 ]    94 ^    95 _
 96 `    97 a    98 b    99 c   100 d   101 e   102 f   103 g
104 h   105 i   106 j   107 k   108 l   109 m   110 n   111 o
112 p   113 q   114 r   115 s   116 t   117 u   118 v   119 w
120 x   121 y   122 z   123 {   124 |   125 }   126 ~

The first entry looks broken and is not. Code 32 is the space character, so it prints a space. Change the %c to %x and you get the table in hexadecimal instead.

Run in Compiler

Where this is used

  • Every text file on your machine. A .c file, a .csv file and this page's HTML are all sequences of these codes, with UTF-8 handling anything above 127.
  • HTTP. A request line such as GET /index.html HTTP/1.1 is ASCII text, ended by CR and LF, codes 13 and 10. That is why a browser and a server written in different languages can talk.
  • Your keyboard. Pressing Ctrl and C sends code 3, ETX, which is why that key combination has meant "stop" since long before your laptop existed.
  • Sorting names. Comparing text compares these codes, and 'Z' is 90 while 'a' is 97. So a plain sort puts every capital before every small letter, which is why real software sorts with a function that ignores case.

Common mistakes

1. Using a digit character as its value.

char digit = '7';
printf("%d\n", digit);        /* prints 55, not 7 */

There is no message. GCC is right that a char is a small integer, and 55 is the number you asked for. Subtract '0' whenever you want the value.

2. Doing case arithmetic on something that is not a letter.

char d = '5';
printf("[%c]\n", d - 32);     /* prints [ ] with an invisible control char */

Again no message, at any warning level. 53 minus 32 is 21, which is NAK, invisible. Module 5 gives you the if that checks the character first.

3. Putting a Bangla letter in a char.

char c = 'āĻĸ';
printf("%d\n", c);

GCC 12 gives you two warnings: warning: multi-character character constant and warning: overflow in conversion from 'int' to 'char' changes value from '14722722' to '-94'. It then prints -94. A char is one byte and that letter needs three.

4. Storing a number a char cannot hold.

char c = 300;
printf("%d\n", c);            /* prints 44 */

GCC 12 says warning: overflow in conversion from 'int' to 'char' changes value from '300' to '44', and builds it anyway. A char on this machine holds -128 to 127, which is exactly why ASCII stops at 127.

Brain teaser

Kenji says characters are just numbers, so adding two of them must be fine. Zara is not convinced.

printf("%d\n", 'a' + 'b');
printf("[%c]\n", 'a' + 'b');

Answer three questions. What does the first line print? What does the second line print, and why is that not a letter? And what is the largest pair of printable ASCII characters whose sum is still a printable ASCII character?

Look up both codes in the table and add them on paper. Then ask what the highest code in the table is, and what a byte does with a number above it.

Exercise 1Easy

Maria is building a tool that shows the code of any key. Read one character and print its decimal code.

Input. One line holding exactly one character, which may be a space.

Output. One line with its decimal ASCII code.

Constraints. The character is printable ASCII, code 32 to 126.

Sample. Input A gives 65. Input a single space gives 32.

#include <stdio.h>

int main(void)
{
    char ch;
    scanf("%c", &ch);

    /* One printf. Which specifier do you need? */

    return 0;
}

Note the scanf("%c", &ch) with no space before the %c. That is deliberate: a space there would skip the space character, and the space is one of the test cases.

Graded against hidden tests in the module's Problems lesson, as char-code.

Run in Compiler
Exercise 2Easy

Change the case of one letter using arithmetic only. No library function, no if.

Input. One line with one capital letter, 'A' to 'Z'.

Output. One line with the same letter in small case.

Constraints. The input is always a capital letter, so you need no check.

Sample. Input M gives m.

#include <stdio.h>

int main(void)
{
    char ch;
    scanf("%c", &ch);

    /* Trick one. Write the 32 as a difference of two characters if you can. */

    return 0;
}

Check yourself. Try it with A and with Z, the two ends of the range. If both work, every letter between them works, because the capitals are unbroken in the table.

Run in Compiler
Exercise 3Medium

Amara is checking a form where three digits arrive as characters, with nothing between them. Add up what they stand for.

Input. One line with exactly three digit characters, no spaces.

Output. One line with the sum of the three digit values.

Constraints. Each character is '0' to '9'.

Sample. Input 407 gives 11, because 4 plus 0 plus 7 is 11.

#include <stdio.h>

int main(void)
{
    char a, b, c;
    scanf("%c%c%c", &a, &b, &c);

    /* Trick two, three times. */

    return 0;
}

Check yourself. 000 must give 0 and 999 must give 27. If you get 144 for 000, you added the codes instead of the values.

Graded as digit-sum-three.

Run in Compiler
Exercise 4Hard

ROT13 is a very old way of hiding text: replace each letter with the one 13 places later, wrapping around from z back to a. Do it for one small letter.

Input. One line with one small letter, 'a' to 'z'.

Output. One line with the letter 13 places later, wrapped inside the alphabet.

Constraints. The input is always a small letter.

Sample. Input a gives n. Input z gives m.

The hard part is the wrap. Adding 13 to 'z' leaves the alphabet. One line of arithmetic fixes it, using the remainder operator % and the alphabet position from trick three. The remainder operator is Module 4, and using it early is allowed.

Check yourself. Run it on a, m, n and z. Applying your program twice to the same letter must give the letter back, which is the property that made ROT13 popular.

Common doubts

  • Why 128 codes and not 256?

    ASCII was designed for seven bits, in an era when the eighth was used for error checking. The free half above 127 was filled differently by every vendor, and Unicode replaced that mess.

  • Is a char signed or unsigned?

    The standard leaves it to the compiler, and on most machines it is signed, holding -128 to 127. For codes 0 to 127 it makes no difference, which is all of ASCII.

  • Can I write 65 instead of 'A'?

    Yes, and the program behaves identically. Do not: 'A' says what you mean and survives being read six months later, and it works on machines where the codes differ.

  • Why does the table show hexadecimal too?

    Because tools show you bytes in hex. A memory dump, a network capture and an escape such as \x41 all use base 16. In those places 41 is easier to recognise than 65.

  • Do the three tricks work on Bangla letters?

    No. One Bangla letter is three bytes in UTF-8, so there is no single number to add 32 to. You need a text library, and Module 10 explains what a C string can and cannot do.

Key takeaways

  • A char is a small integer; %c and %d are two lenses on the same byte.
  • Four anchors carry the table: 48, 65, 97 and 32.
  • Capital to small is plus 32; digit character to value is minus '0'; letter to position is minus 'A'.
  • Codes 0 to 31 are control characters, and five of them are still in daily use.
  • The digits, the capitals and the small letters each sit in one unbroken run.
  • ASCII stops at 127; one Bangla letter is three UTF-8 bytes, so the tricks do not reach it.

Next you will write code another human can read. It is the one skill in this module the compiler will never check for you.

End of lesson 4

Mark it done, and your progress moves with you.

Next: Comments, Whitespace and Code a Human Can Read