Module 2 ¡ Variables and Data Types
Integer Types, Their Ranges, and Overflow
In this lesson
- Choose between
char,short,int,long,long longand their unsigned forms. - State what the C standard promises about a type's size, and what it refuses to promise.
- Predict what happens past the top of a range: unsigned wraps, signed is undefined.
Kenji builds a step counter. It works for two years, and then one user's total goes negative overnight. That user had walked past 2,147,483,647 steps.
Nothing was broken. The box was full, and C did what a full box does. This lesson is about where the top of a box is, and what is on the other side of it.
The integer family
An integer is a whole number, with no decimal point. C gives you five sizes of box, each in a signed and an unsigned flavour.
| Type | Size on the Playground | Range on the Playground | Specifier |
|---|---|---|---|
signed char | 1 byte | -128 to 127 | %d |
short | 2 bytes | -32,768 to 32,767 | %hd |
int | 4 bytes | -2,147,483,648 to 2,147,483,647 | %d |
long | 8 bytes | about -9.2 to 9.2 quintillion | %ld |
long long | 8 bytes | about -9.2 to 9.2 quintillion | %lld |
unsigned char | 1 byte | 0 to 255 | %u |
unsigned short | 2 bytes | 0 to 65,535 | %hu |
unsigned int | 4 bytes | 0 to 4,294,967,295 | %u |
unsigned long long | 8 bytes | 0 to about 18.4 quintillion | %llu |
Read the heading of that table again. It says "on the Playground", and that is not politeness.
What the standard promises, and what it does not
The C standard sets minimum ranges, not exact sizes. It promises this and no more.
charis 1 byte, always, by definition. A byte is at least 8 bits.shortandintreach at least -32,767 to 32,767.longreaches at least about plus or minus 2.1 billion.long longreaches at least about plus or minus 9.2 quintillion.- Each type in that list is at least as wide as the one before it.
So int is 4 bytes on every desktop and phone you will meet. It is 2 bytes on small chips inside washing machines and car dashboards, and those chips run C too.
long is the one to distrust. It is 8 bytes on Linux and macOS, and 4 bytes on Windows, on the same processor. Lesson 5 shows you how to ask instead of guessing.
So the honest habit is: int for ordinary counting, long long when a number might pass 2 billion. And never write a size into your head as a fact.
Signed and unsigned: the same bits, two readings
Every integer type comes in two flavours. A signed type spends one bit on the sign, so half its range is negative. An unsigned type has no negatives and spends that bit on size instead, which doubles the top.
#include <stdio.h>
int main(void)
{
int temperature = -12; /* signed by default */
unsigned int student_count = 350; /* never negative, so unsigned fits */
printf("temperature : %d\n", temperature);
printf("student count : %u\n", student_count);
return 0;
}
temperature : -12
student count : 350
Plain int means signed int. You write unsigned when you want the other reading.
Use unsigned when a negative value would be meaningless: a count of people, a size in bytes, a position in a list. Then read the next two sections before you trust it.
The limits live in a header
You do not memorise the ranges. You print them, from <limits.h>, on the machine you are actually running on.
#include <stdio.h>
#include <limits.h>
int main(void)
{
printf("signed char : %d to %d\n", SCHAR_MIN, SCHAR_MAX);
printf("short : %d to %d\n", SHRT_MIN, SHRT_MAX);
printf("int : %d to %d\n", INT_MIN, INT_MAX);
printf("long long : %lld to %lld\n", LLONG_MIN, LLONG_MAX);
printf("unsigned int: 0 to %u\n", UINT_MAX);
return 0;
}
signed char : -128 to 127
short : -32768 to 32767
int : -2147483648 to 2147483647
long long : -9223372036854775808 to 9223372036854775807
unsigned int: 0 to 4294967295
long is missing from that list on purpose. Its answer is 8 on the Playground and 4 on a Windows machine. A printed list of ranges cannot contain it and stay true everywhere.
So the header is the answer to "how big is it", and lesson 5 turns that into a habit.
Unsigned wraps, and the standard says so
Unsigned arithmetic is defined to wrap, like a car odometer rolling from 999999 back to 000000. This is a rule, not an accident.
#include <stdio.h>
int main(void)
{
unsigned int stock = 0;
printf("stock : %u\n", stock);
stock = stock - 1;
printf("stock - 1 : %u\n", stock);
unsigned int big = 4294967295u;
printf("big + 1 : %u\n", big + 1u);
return 0;
}
stock : 0
stock - 1 : 4294967295
big + 1 : 0
Nothing went wrong there. Every one of those three numbers is promised by the standard, on every machine, for ever.
That is also why it bites. A shop that has 0 items in stock and sells 1 now has four billion items, and no message was printed anywhere.
So unsigned is right for a value that genuinely cannot be negative, and wrong for a value you are going to subtract from.
Signed overflow is undefined, which is worse than wrong
Push a signed integer past its top and the C standard stops describing your program. Not "gives a strange number". Stops describing it.
That phrase is undefined behaviour, and it means the compiler is allowed to assume the situation never happens. It then builds your program on that assumption.
On the machine we ran this on, the value wrapped to the bottom of the range, which is what most people expect. The same program also answered a question about itself incorrectly, in the same run, because the compiler had assumed the wrap away.
Both of those are allowed. Neither is a bug in the compiler. The bug is in the program that overflowed.
So the habit is: when a calculation might pass about 2 billion, use long long before it does, not after.
Suffixes: telling the compiler what kind of number you wrote
A plain number in your code, like 42, is an int. Add a letter and you change what it is.
Literal suffixes and their specifiers
42 int printf("%d", 42)
42u unsigned int printf("%u", 42u)
42L long printf("%ld", 42L)
42LL long long printf("%lld", 42LL)
42ULL unsigned long long printf("%llu", 42ULL)
- Capital
Lis easier to read than smalll, which looks like a 1. 2147483647 + 1overflows, because both sides areint.2147483647LL + 1does not, because the left side is already wide.- The specifier has to match the type, or
printfreads the wrong bytes.
So the suffix is how you widen a calculation before it starts, and the specifier is how you read the answer out.
The smallest possible demonstration: the largest int, then the same sum done in a wider box.
#include <stdio.h>
#include <limits.h>
int main(void)
{
int top = INT_MAX;
long long roomy = (long long)INT_MAX + 1;
printf("int can reach : %d\n", top);
printf("long long holds : %lld\n", roomy);
return 0;
}
int can reach : 2147483647
long long holds : 2147483648
The (long long) in front is a cast, and Module 4 explains it. It widens the value before the + 1 happens, which is the whole trick.
Unsigned wrap, from both ends, with every number promised by the standard.
#include <stdio.h>
int main(void)
{
unsigned int stock = 0;
unsigned int big = 4294967295u;
printf("0 - 1 as unsigned : %u\n", stock - 1u);
printf("top + 1 : %u\n", big + 1u);
printf("top + 2 : %u\n", big + 2u);
return 0;
}
0 - 1 as unsigned : 4294967295
top + 1 : 0
top + 2 : 1
Count the third line. It is 1, not 2, because the wrap happened once and then the counting continued normally.
Run in CompilerSteps per day times days. Two boxes, two answers, one input.
#include <stdio.h>
int main(void)
{
int steps_per_day = 0;
int days = 0;
scanf("%d %d", &steps_per_day, &days);
int as_int = steps_per_day * days;
long long as_long_long = (long long)steps_per_day * days;
printf("as int : %d\n", as_int);
printf("as long long : %lld\n", as_long_long);
return 0;
}
as int : -294967296
as long long : 4000000000
That output is for the input 20000 200000. Only the second line is a promise. The first is what this run happened to print, because the multiplication overflowed a signed type and the standard then says nothing.
Look at where the cast sits. It is on the left operand, before the multiplication, so the whole sum is done in the wide type. Writing (long long)(steps_per_day * days) would overflow first and widen the wreckage afterwards.
Where this is used
- The year 2038 problem. Unix has stored time as a signed 32-bit count of seconds since 1970. That counter fills on 19 January 2038, so Linux moved
time_tto 64 bits, and embedded devices that did not are still being found. - File sizes. Linux uses a 64-bit
off_tso a file may pass 2 GB. The 32-bit version is still there under the nameoff_ton old builds, which is why a 2 GB limit used to appear everywhere at once. - IPv4 addresses. An address is a 32-bit unsigned integer, which is why there are 4,294,967,296 of them and why the world ran out.
- Colours. A pixel's red, green and blue channels are three
unsigned charvalues, 0 to 255 each. That range is the type, not a convention. - Ariane 5, flight 501, 1996. A 64-bit floating point value was converted into a 16-bit signed integer that could not hold it. The rocket was lost 37 seconds after launch. The code was correct on the previous rocket, which flew slower.
Common mistakes
1. A literal that does not fit the box.
int big = 3000000000;
printf("%d\n", big);
GCC 12 prints no message, and neither does a local gcc -Wall -Wextra. The program builds and prints -1294967296. Three billion does not fit in a 4-byte signed box, and the value that arrives is not the one you typed. Use long long.
2. The wrong specifier for a wide value.
long long big = 5000000000LL;
printf("%d\n", big);
The Playground compiles this silently, because the format check lives behind -Wall. It printed 705032704 on our run. %d told printf to read 4 bytes, and the value is 8 bytes long. Write %lld.
3. An overflow check that the compiler is allowed to delete.
int x;
scanf("%d", &x);
printf("did it overflow? %d\n", x + 1 < x);
No message, at either command line. With the input 2147483647 this printed 0, meaning "no overflow", on a run where the stored value of x + 1 was -2147483648. Signed overflow is undefined, so the compiler may assume x + 1 is always larger than x and answer without looking. Check the range before you add, never after.
4. Comparing a signed value with an unsigned one.
int a = -1;
unsigned int b = 1;
printf("%d\n", a < b);
The Playground says nothing; a local gcc -Wextra says warning: comparison of integer expressions of different signedness. It prints 0, because -1 is converted to unsigned first and becomes 4294967295. Keep both sides of a comparison in the same family.
Amara adds up two readings from a sensor. Each one fits in an int, and their sum does not.
Input. One line with two integers a b.
Output. One line with their sum.
Constraints. 0 <= a, b <= 2000000000.
Sample. Input 2000000000 2000000000 gives 4000000000.
#include <stdio.h>
int main(void)
{
int a = 0;
int b = 0;
scanf("%d %d", &a, &b);
/* The inputs fit in an int. The answer does not. */
return 0;
}
Graded as safe-sum. Both ends of the constraints are hidden tests, and one of them is the whole problem.
A server has been up for a very large number of seconds. Print that as days, hours, minutes and seconds.
Input. One line with one integer s, the number of seconds.
Output. One line: d days, h hours, m minutes, s seconds, with the four numbers filled in.
Constraints. 0 <= s <= 1000000000000.
Sample. Input 90061 gives 1 days, 1 hours, 1 minutes, 1 seconds. The grammar is deliberately not fixed; print it exactly as shown.
#include <stdio.h>
int main(void)
{
long long s = 0;
scanf("%lld", &s);
/* 86400 seconds in a day, 3600 in an hour, 60 in a minute. */
return 0;
}
Graded as seconds-to-days. Dividing one whole number by another throws the fraction away, which is exactly what you want here.
Predict before you run. Read two unsigned values and print their difference as unsigned, then the same difference done in a signed box.
Input. One line with two non-negative integers a b, each at most 4294967295.
Output. Two lines: the unsigned difference, then the signed one.
Constraints. Try 3 5 and 0 1 by hand first.
#include <stdio.h>
int main(void)
{
unsigned int a = 0;
unsigned int b = 0;
scanf("%u %u", &a, &b);
printf("%u\n", a - b);
printf("%lld\n", (long long)a - (long long)b);
return 0;
}
Not graded in this module. Run it with 3 5 and write down both numbers, then explain the first one to somebody in one sentence.
A student identity number is too long for an int. Split it into its top part and its bottom nine digits, using nothing but division and subtraction.
Input. One line with one integer id, between 1000000000 and 999999999999999999.
Output. Two lines: id / 1000000000, then the part that division threw away.
Constraints. The value never fits in an int, so every variable here is a long long.
Sample. Input 1234567890123 gives 1234 then 567890123.
#include <stdio.h>
int main(void)
{
long long id = 0;
scanf("%lld", &id);
long long top = id / 1000000000LL;
long long rest = 0; /* what is left when the top is taken away */
printf("%lld\n", top);
printf("%lld\n", rest);
return 0;
}
Not graded in this module. The remainder operator % arrives in Module 4 and would do this in one symbol; getting there with - and * first is the point.
Common doubts
Should I just use
long longeverywhere and stop worrying?For a beginner's programs, almost. It costs 4 extra bytes per value and removes a whole class of bug. In a large array or a network packet, those bytes start to matter, and then you choose.
Is
long longslower thanint?Not measurably on a 64-bit machine, which is every desktop and phone. On a small 8-bit chip it is much slower, because the processor has to do the sum in pieces.
What is the difference between
longandlong long?On the Playground, nothing: both are 8 bytes. On Windows,
longis 4 bytes andlong longis 8. That is exactly why this track useslong longand notlong.Why is unsigned wrap defined but signed overflow not?
Unsigned arithmetic is defined as counting in a circle, which every machine does the same way. Signed formats once differed between machines, and leaving it undefined also lets compilers optimise more aggressively.
How do I tell whether my answer overflowed?
By checking the inputs before you calculate, against
INT_MAXorLLONG_MAX. Checking afterwards asks a question the language has already refused to answer.
Key takeaways
- C promises minimum ranges and an ordering, never an exact size for any type but
char. intis 4 bytes on everything you will meet, and 2 bytes on chips that still exist.longis 8 bytes on Linux and 4 on Windows, so this track writeslong long.- Unsigned arithmetic wraps, by the standard's own rule, and 0 minus 1 is a huge number.
- Signed overflow is undefined: the compiler may assume it never happens and act on that.
- Widen with a suffix or a cast before the calculation, and match the specifier to the type.
Next are the numbers with a decimal point, where the surprise is not the size of the box but the fact that 0.1 was never quite 0.1.
End of lesson 2
Mark it done, and your progress moves with you.
Next: Floating Point: float, double, and Why 0.1 Is Not 0.1