Module 1 ¡ Meet C: Your First Programs, Symbol by Symbol
The Complete ASCII Table, All 128 Rows
In this lesson
- Explain why a
charin C is really a small integer. - Recall the four anchors:
'0'is 48,'A'is 65,'a'is 97, a space is 32. - Use the three arithmetic tricks, and read any row of the complete table below.
A computer stores numbers. Only numbers. So how does it store the letter A?
It agrees on a code. In 1963 a standard wrote the agreement down: A is 65, B is 66, and so on to 127. Its name is the American Standard Code for Information Interchange, ASCII for short. Every machine that follows it reads the same byte the same way.
This lesson is a reference page with the whole agreement printed in it, and five short programs that use it.
One byte, two readings
In C, a char is not really a letter. It is a small integer that you have agreed to show on the screen as a letter.
That is why the same variable can be printed two ways. %c says "show me the character this number stands for". %d says "show me the number".
#include <stdio.h>
int main(void)
{
char letter = 'A';
printf("As a character: %c\n", letter);
printf("As a number: %d\n", letter);
return 0;
}
As a character: A
As a number: 65
Nothing was converted between those two lines. The byte in memory never changed. Only the lens you looked at it through changed.
The four anchors and the three tricks
You do not need to learn 128 numbers. You need four, and everything else follows by arithmetic.
| Character | Code | The whole group |
|---|---|---|
'0' | 48 | The digits '0' to '9' are 48 to 57 |
'A' | 65 | The capitals 'A' to 'Z' are 65 to 90 |
'a' | 97 | The small letters 'a' to 'z' are 97 to 122 |
| a space | 32 | The first printable character of all |
From those four, three tricks cover most of the real work you will do with characters.
- Change case by 32. A small letter is exactly 32 above its capital, so
'a' - 'A'is 32. - Digit character to digit value: subtract
'0'.'7' - '0'is 7, because 55 minus 48 is 7. - Letter to alphabet position: subtract
'A'or'a'.'D' - 'A'is 3, so D is the fourth letter.
Notice what all three have in common. You never write 48, 65 or 97 in your code. You write '0', 'A' or 'a', and the compiler fills the number in, which keeps the line readable.
So if you can subtract, you can do case conversion, digit parsing and alphabet positions without a single library function.
The control characters: 0 to 31, and 127
The first 32 codes are not printable. They are control characters: commands left over from teleprinters, machines that typed text down a wire. Most are museum pieces. Five still matter, and they are in bold.
| Dec | Hex | Name | Full name | What it does |
|---|---|---|---|---|
| 0 | 00 | NUL | Null | Ends a C string. The one you will meet in Module 10 |
| 1 | 01 | SOH | Start of heading | A teleprinter header marker, unused today |
| 2 | 02 | STX | Start of text | Marked where the message began |
| 3 | 03 | ETX | End of text | What Ctrl and C send to a terminal |
| 4 | 04 | EOT | End of transmission | What Ctrl and D send, read as end of input |
| 5 | 05 | ENQ | Enquiry | Asked the other machine to answer |
| 6 | 06 | ACK | Acknowledge | The other machine's yes |
| 7 | 07 | BEL | Bell | Rings the terminal bell. Written \a |
| 8 | 08 | BS | Backspace | Moves back one character. Written \b |
| 9 | 09 | HT | Horizontal tab | Jumps to the next tab stop. Written \t |
| 10 | 0A | LF | Line feed | The newline on Linux and macOS. Written \n |
| 11 | 0B | VT | Vertical tab | Moved down a fixed distance on paper |
| 12 | 0C | FF | Form feed | Ejects a page on a printer. Written \f |
| 13 | 0D | CR | Carriage return | Back to the start of the line. Windows writes CR then LF |
| 14 | 0E | SO | Shift out | Switched to an alternative character set |
| 15 | 0F | SI | Shift in | Switched back again |
| 16 | 10 | DLE | Data link escape | Told the link that what follows is a command |
| 17 | 11 | DC1 | Device control 1 | Also called XON: resume sending |
| 18 | 12 | DC2 | Device control 2 | A spare device signal |
| 19 | 13 | DC3 | Device control 3 | Also called XOFF: stop sending |
| 20 | 14 | DC4 | Device control 4 | A spare device signal |
| 21 | 15 | NAK | Negative acknowledge | The other machine's no |
| 22 | 16 | SYN | Synchronous idle | Filled a quiet line to keep the timing |
| 23 | 17 | ETB | End of transmission block | Ended one block of a longer message |
| 24 | 18 | CAN | Cancel | Said the data before it was wrong |
| 25 | 19 | EM | End of medium | The tape or the card ran out |
| 26 | 1A | SUB | Substitute | Ctrl and Z on Windows, read as end of file |
| 27 | 1B | ESC | Escape | Starts a terminal colour code. Written \033 |
| 28 | 1C | FS | File separator | The coarsest of the four separators |
| 29 | 1D | GS | Group separator | Separated groups inside a file |
| 30 | 1E | RS | Record separator | Separated records inside a group |
| 31 | 1F | US | Unit separator | Separated fields inside a record |
| 127 | 7F | DEL | Delete | All seven holes punched on paper tape, which meant ignore this |
The five in bold are the ones you will actually type. NUL ends a string, HT is \t and LF is \n. CR is the other half of a Windows line ending, and ESC starts the codes that colour a terminal.
So a text file with no control characters in it at all would be one endless line.
The printable half: 32 to 126, complete
Everything from 32 to 126 has a shape you can see. Read the table in four blocks, left to right. The first cell of the first block is the space character, which is why it looks empty.
| Dec | Hex | Char | Dec | Hex | Char | Dec | Hex | Char | Dec | Hex | Char |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 32 | 20 | space | 56 | 38 | 8 | 80 | 50 | P | 104 | 68 | h |
| 33 | 21 | ! | 57 | 39 | 9 | 81 | 51 | Q | 105 | 69 | i |
| 34 | 22 | " | 58 | 3A | : | 82 | 52 | R | 106 | 6A | j |
| 35 | 23 | # | 59 | 3B | ; | 83 | 53 | S | 107 | 6B | k |
| 36 | 24 | $ | 60 | 3C | < | 84 | 54 | T | 108 | 6C | l |
| 37 | 25 | % | 61 | 3D | = | 85 | 55 | U | 109 | 6D | m |
| 38 | 26 | & | 62 | 3E | > | 86 | 56 | V | 110 | 6E | n |
| 39 | 27 | ' | 63 | 3F | ? | 87 | 57 | W | 111 | 6F | o |
| 40 | 28 | ( | 64 | 40 | @ | 88 | 58 | X | 112 | 70 | p |
| 41 | 29 | ) | 65 | 41 | A | 89 | 59 | Y | 113 | 71 | q |
| 42 | 2A | * | 66 | 42 | B | 90 | 5A | Z | 114 | 72 | r |
| 43 | 2B | + | 67 | 43 | C | 91 | 5B | [ | 115 | 73 | s |
| 44 | 2C | , | 68 | 44 | D | 92 | 5C | \ | 116 | 74 | t |
| 45 | 2D | - | 69 | 45 | E | 93 | 5D | ] | 117 | 75 | u |
| 46 | 2E | . | 70 | 46 | F | 94 | 5E | ^ | 118 | 76 | v |
| 47 | 2F | / | 71 | 47 | G | 95 | 5F | _ | 119 | 77 | w |
| 48 | 30 | 0 | 72 | 48 | H | 96 | 60 | ` | 120 | 78 | x |
| 49 | 31 | 1 | 73 | 49 | I | 97 | 61 | a | 121 | 79 | y |
| 50 | 32 | 2 | 74 | 4A | J | 98 | 62 | b | 122 | 7A | z |
| 51 | 33 | 3 | 75 | 4B | K | 99 | 63 | c | 123 | 7B | { |
| 52 | 34 | 4 | 76 | 4C | L | 100 | 64 | d | 124 | 7C | | |
| 53 | 35 | 5 | 77 | 4D | M | 101 | 65 | e | 125 | 7D | } |
| 54 | 36 | 6 | 78 | 4E | N | 102 | 66 | f | 126 | 7E | ~ |
| 55 | 37 | 7 | 79 | 4F | O | 103 | 67 | g |
Three patterns are worth seeing in the table itself. The digits sit together, the capitals sit together, and the small letters sit together, each group unbroken.
Look at 65 and 97. The gap between a capital and its small letter is 32 on every row, which is trick one, visible.
Look at 48 and 57. The digits are in order with no gaps, which is what makes trick two work for every digit and not just for '7'.
Beyond ASCII, and what that means for Bangla
ASCII only defines 0 to 127, because it was designed for seven bits. A byte holds eight, so codes 128 to 255 were left free and every vendor filled them differently. That mess is called "extended ASCII" and it is why the same file looked like nonsense on a different machine.
Unicode fixed it by giving every character in every script its own number, and UTF-8 is the way those numbers are stored as bytes. ASCII is the first 128 of them, unchanged, which is why old files still work.
Here is the part that matters for you. In UTF-8, one Bangla letter takes three bytes, not one. The word "āĻĸāĻžāĻāĻž" is four letters and twelve bytes, and strlen answers 12.
So a char holds one byte, not one letter. The three tricks work on ASCII letters and digits. They do not work on Bangla text, because there is no single number to add 32 to. Real Bangla text handling is a Module 10 topic and needs a library.
The smallest program that shows the whole idea. One variable, four printed lines, no conversion anywhere.
#include <stdio.h>
int main(void)
{
char letter = 'A';
printf("As a character: %c\n", letter);
printf("As a number: %d\n", letter);
printf("Add 1: %c\n", letter + 1);
printf("Add 32: %c\n", letter + 32);
return 0;
}
As a character: A
As a number: 65
Add 1: B
Add 32: a
Adding 1 moved along the alphabet. Adding 32 jumped from the capitals to the small letters. Both are ordinary arithmetic on the number 65.
Run in CompilerTrick one, in both directions. Add 32 to go down to small letters, subtract 32 to go up to capitals.
#include <stdio.h>
int main(void)
{
char upper = 'M';
char lower = 'q';
/* A small letter is exactly 32 above its capital. */
printf("%c becomes %c\n", upper, upper + 32);
printf("%c becomes %c\n", lower, lower - 32);
return 0;
}
M becomes m
q becomes Q
This only works if the character really is a letter. Try it on '5' and you get code 21, an invisible control character, with no warning at all.
Trick two, and the mistake it exists to prevent. The character '7' is the number 55, and the number 7 is something you have to work out.
#include <stdio.h>
int main(void)
{
char digit = '7';
int wrong = digit; /* 55, the ASCII code */
int right = digit - '0'; /* 7, the value it stands for */
printf("Wrong: %d\n", wrong);
printf("Right: %d\n", right);
printf("Squared: %d\n", right * right);
return 0;
}
Wrong: 55
Right: 7
Squared: 49
That - '0' appears in almost every program that reads digits out of text, which is most programs that read anything.
This one uses a loop, which is Module 6. Read it, run it, and do not worry about writing your own yet. The point is that the table above is not a fact to trust, it is something your machine can print for you.
#include <stdio.h>
int main(void)
{
for (int code = 32; code <= 126; code++) {
printf("%3d %c", code, code);
/* Start a new line after every eighth entry. */
if ((code - 31) % 8 == 0) {
printf("\n");
} else {
printf(" ");
}
}
printf("\n");
return 0;
}
32 33 ! 34 " 35 # 36 $ 37 % 38 & 39 '
40 ( 41 ) 42 * 43 + 44 , 45 - 46 . 47 /
48 0 49 1 50 2 51 3 52 4 53 5 54 6 55 7
56 8 57 9 58 : 59 ; 60 < 61 = 62 > 63 ?
64 @ 65 A 66 B 67 C 68 D 69 E 70 F 71 G
72 H 73 I 74 J 75 K 76 L 77 M 78 N 79 O
80 P 81 Q 82 R 83 S 84 T 85 U 86 V 87 W
88 X 89 Y 90 Z 91 [ 92 \ 93 ] 94 ^ 95 _
96 ` 97 a 98 b 99 c 100 d 101 e 102 f 103 g
104 h 105 i 106 j 107 k 108 l 109 m 110 n 111 o
112 p 113 q 114 r 115 s 116 t 117 u 118 v 119 w
120 x 121 y 122 z 123 { 124 | 125 } 126 ~
The first entry looks broken and is not. Code 32 is the space character, so it prints a space. Change the %c to %x and you get the table in hexadecimal instead.
Where this is used
- Every text file on your machine. A
.cfile, a.csvfile and this page's HTML are all sequences of these codes, with UTF-8 handling anything above 127. - HTTP. A request line such as
GET /index.html HTTP/1.1is ASCII text, ended byCRandLF, codes 13 and 10. That is why a browser and a server written in different languages can talk. - Your keyboard. Pressing Ctrl and C sends code 3,
ETX, which is why that key combination has meant "stop" since long before your laptop existed. - Sorting names. Comparing text compares these codes, and
'Z'is 90 while'a'is 97. So a plain sort puts every capital before every small letter, which is why real software sorts with a function that ignores case.
Common mistakes
1. Using a digit character as its value.
char digit = '7';
printf("%d\n", digit); /* prints 55, not 7 */
There is no message. GCC is right that a char is a small integer, and 55 is the number you asked for. Subtract '0' whenever you want the value.
2. Doing case arithmetic on something that is not a letter.
char d = '5';
printf("[%c]\n", d - 32); /* prints [ ] with an invisible control char */
Again no message, at any warning level. 53 minus 32 is 21, which is NAK, invisible. Module 5 gives you the if that checks the character first.
3. Putting a Bangla letter in a char.
char c = 'āĻĸ';
printf("%d\n", c);
GCC 12 gives you two warnings: warning: multi-character character constant and warning: overflow in conversion from 'int' to 'char' changes value from '14722722' to '-94'. It then prints -94. A char is one byte and that letter needs three.
4. Storing a number a char cannot hold.
char c = 300;
printf("%d\n", c); /* prints 44 */
GCC 12 says warning: overflow in conversion from 'int' to 'char' changes value from '300' to '44', and builds it anyway. A char on this machine holds -128 to 127, which is exactly why ASCII stops at 127.
Maria is building a tool that shows the code of any key. Read one character and print its decimal code.
Input. One line holding exactly one character, which may be a space.
Output. One line with its decimal ASCII code.
Constraints. The character is printable ASCII, code 32 to 126.
Sample. Input A gives 65. Input a single space gives 32.
#include <stdio.h>
int main(void)
{
char ch;
scanf("%c", &ch);
/* One printf. Which specifier do you need? */
return 0;
}
Note the scanf("%c", &ch) with no space before the %c. That is deliberate: a space there would skip the space character, and the space is one of the test cases.
Graded against hidden tests in the module's Problems lesson, as char-code.
Change the case of one letter using arithmetic only. No library function, no if.
Input. One line with one capital letter, 'A' to 'Z'.
Output. One line with the same letter in small case.
Constraints. The input is always a capital letter, so you need no check.
Sample. Input M gives m.
#include <stdio.h>
int main(void)
{
char ch;
scanf("%c", &ch);
/* Trick one. Write the 32 as a difference of two characters if you can. */
return 0;
}
Check yourself. Try it with A and with Z, the two ends of the range. If both work, every letter between them works, because the capitals are unbroken in the table.
Amara is checking a form where three digits arrive as characters, with nothing between them. Add up what they stand for.
Input. One line with exactly three digit characters, no spaces.
Output. One line with the sum of the three digit values.
Constraints. Each character is '0' to '9'.
Sample. Input 407 gives 11, because 4 plus 0 plus 7 is 11.
#include <stdio.h>
int main(void)
{
char a, b, c;
scanf("%c%c%c", &a, &b, &c);
/* Trick two, three times. */
return 0;
}
Check yourself. 000 must give 0 and 999 must give 27. If you get 144 for 000, you added the codes instead of the values.
Graded as digit-sum-three.
ROT13 is a very old way of hiding text: replace each letter with the one 13 places later, wrapping around from z back to a. Do it for one small letter.
Input. One line with one small letter, 'a' to 'z'.
Output. One line with the letter 13 places later, wrapped inside the alphabet.
Constraints. The input is always a small letter.
Sample. Input a gives n. Input z gives m.
The hard part is the wrap. Adding 13 to 'z' leaves the alphabet. One line of arithmetic fixes it, using the remainder operator % and the alphabet position from trick three. The remainder operator is Module 4, and using it early is allowed.
Check yourself. Run it on a, m, n and z. Applying your program twice to the same letter must give the letter back, which is the property that made ROT13 popular.
Common doubts
Why 128 codes and not 256?
ASCII was designed for seven bits, in an era when the eighth was used for error checking. The free half above 127 was filled differently by every vendor, and Unicode replaced that mess.
Is a
charsigned or unsigned?The standard leaves it to the compiler, and on most machines it is signed, holding -128 to 127. For codes 0 to 127 it makes no difference, which is all of ASCII.
Can I write 65 instead of
'A'?Yes, and the program behaves identically. Do not:
'A'says what you mean and survives being read six months later, and it works on machines where the codes differ.Why does the table show hexadecimal too?
Because tools show you bytes in hex. A memory dump, a network capture and an escape such as
\x41all use base 16. In those places 41 is easier to recognise than 65.Do the three tricks work on Bangla letters?
No. One Bangla letter is three bytes in UTF-8, so there is no single number to add 32 to. You need a text library, and Module 10 explains what a C string can and cannot do.
Key takeaways
- A
charis a small integer;%cand%dare two lenses on the same byte. - Four anchors carry the table: 48, 65, 97 and 32.
- Capital to small is plus 32; digit character to value is minus
'0'; letter to position is minus'A'. - Codes 0 to 31 are control characters, and five of them are still in daily use.
- The digits, the capitals and the small letters each sit in one unbroken run.
- ASCII stops at 127; one Bangla letter is three UTF-8 bytes, so the tricks do not reach it.
Next you will write code another human can read. It is the one skill in this module the compiler will never check for you.
End of lesson 4
Mark it done, and your progress moves with you.
Next: Comments, Whitespace and Code a Human Can Read