Learn C++ STL

Lesson 1 of 9 · string: Text That Knows Its Own Length

Module 3 · string: Text That Knows Its Own Length

string: A Row of Characters That Knows Its Own Length

FreeReading

In this lesson

  • Declare a string, read a word with cin >> and a whole line with getline, and print its size.
  • Read and change one character by index, join strings with +, and compare them with ==.
  • Explain why a Bangla name's size() is not the number of letters you see, using its bytes.

Maria types her name into a C program that has char name[10]. "Maria" fits. Her full name, "Maria Lopez", does not, and the extra letters spill into memory the array does not own. Every strlen then walks the letters again to count them. In C++ none of that is your job: string name; grows to fit, and name.size() already knows the length. This lesson shows the string, and one honest surprise about Maria's name in Bangla.

The problem: text that will not sit still

In the C track, text was a char array ending in '\0'. You chose its length before the program ran. You copied with strcpy, joined with strcat and measured with strlen, which counts the characters one by one every time you call it. A string is the C++ answer: a row of characters that grows by itself and keeps its own length. Its full name is std::string, and it lives in the header <string>.

Here is the same job twice. Read a first name and a last name, join them with a space, and print the result and its length. First the C way, which still compiles as C++.

#include <cstdio>
#include <cstring>

int main() {
    char first[20], last[20], full[41];
    scanf("%19s %19s", first, last);
    strcpy(full, first);
    strcat(full, " ");
    strcat(full, last);
    printf("%s has %zu characters\n", full, strlen(full));
    return 0;
}
Maria Lopez has 11 characters

Now the C++ way, on the same input.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string first, last;
    cin >> first >> last;
    string full = first + " " + last;
    cout << full << " has " << full.size() << " characters\n";
    return 0;
}
Maria Lopez has 11 characters

That output is for the input Maria Lopez. The C version needs three array sizes, a %19s to stop an overflow, and two library calls to join. The C++ version has no size anywhere. So a string is text that you never size by hand.

A string is a row of characters, and a little more

Inside, a string keeps its characters side by side in one block of memory, like a C array or a vector from Module 2. A character here is one char, one byte. The string also keeps two numbers: its size, how many characters it holds, and its capacity, how many fit before it needs a bigger block.

One more thing sits in the block. After the last character, the library keeps a '\0', the same end marker C strings use. It is not counted in the size. It is there so the string can hand its text to a C function in one step (lesson 02 shows c_str()).

A string: a small handle, its characters, and the end marker after them string name = "Maria"; name (the handle) where it starts size: 5 capacity: 15 size: 5 characters M a r i a \0 spare room [0] [1] [2] [3] [4] [5] end marker, not in size capacity: 15 characters fit before a bigger block is needed size() is stored, so asking for it costs nothing. strlen in C counted the characters every time.

So a string is a vector of characters that also keeps an end marker, which is why name.size() is instant where strlen(name) was a walk.

The syntax you need on day one

A string and its everyday calls

#include <string>

string s;                     an empty string, ""
string s = "Maria";           a copy of this text
string s(5, '*');             5 copies of one character, "*****"
s.size()   s.length()         how many characters (bytes) it holds
s.empty()                     true when it holds none
s[i]                          the character at index i, 0 to s.size() - 1
s + t      s += t             a new joined string; add t to the end of s
s == t                        true when the two texts are the same
cin >> s;                     read one word, up to a space or a line break
getline(cin, s);              read a whole line, spaces included
  • string: the type. With using namespace std; you write string; without it, std::string.
  • size() and length(): two names for the same number. Containers say size(), so this track does too.
  • s[i]: one char, which you can read or change. Like a vector, nothing checks that i is in range.
  • "Maria" in double quotes is text; 'M' in single quotes is one char. A string is built from either.

So a string is declared like any variable, and everything you do to it is an operator or a call with a dot.

Four ways to make one

You will make strings in four shapes. This program makes one of each and prints it between brackets, so an empty one still shows.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string empty_one;
    string name = "Maria";
    string stars(5, '*');
    string copy = name;
    copy[0] = 'm';

    cout << "[" << empty_one << "] size " << empty_one.size() << '\n';
    cout << "[" << name << "] size " << name.size() << '\n';
    cout << "[" << stars << "] size " << stars.size() << '\n';
    cout << "[" << copy << "] size " << copy.size() << '\n';
    return 0;
}
[] size 0
[Maria] size 5
[*****] size 5
[maria] size 5

The table reads the four lines back.

DeclarationWhat you getUse it when
string s;no characters, ""you will read into it or add to it
string s = "text";a copy of the textyou know the text when you write the program
string s(n, c);n copies of the character ca line of dashes, a row of stars, padding
string s = t;a copy of the string tyou want to change one and keep the other

The last line shows something C never did with =. Changing copy[0] left name alone. In C, char* copy = name; makes a second name for the same letters. So = on strings copies the text; it does not share it.

Reading a word, and reading a line

cin >> s skips any spaces, then reads one word: characters up to the next space, tab or line break. getline(cin, s) reads everything up to the end of the line, spaces included, and throws the line break away. Module 1, lesson 02, met both. Here they are on one line of input.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string word, rest;
    cin >> word;
    getline(cin, rest);
    cout << "word: [" << word << "]\n";
    cout << "rest: [" << rest << "]\n";
    cout << "rest has " << rest.size() << " characters\n";
    return 0;
}
word: [Maria]
rest: [ Lopez likes tea]
rest has 16 characters

That output is for the input Maria Lopez likes tea. >> stopped at the first space and left it in the input. getline then took the rest of the line, starting with that space. This is also the trap Module 1 showed: after cin >> n, a getline reads the empty rest of the number's line. So read words with >>, lines with getline, and finish a line before you switch.

Indexing, and Bob's last index

The characters are numbered from 0, like an array's boxes. A string of size n has indexes 0 to n - 1, and s[i] is one char you can read or change. Bob prints the character codes of "Maria", one per index. He writes <= to be sure he reaches the last one.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string name = "Maria";
    for (size_t i = 0; i <= name.size(); i++) {
        cout << "[" << (int)name[i] << "]";
    }
    cout << '\n';
    return 0;
}
[77][97][114][105][97][0]

Five codes are the letters: 77 is 'M' and 97 is 'a', from the ASCII table. The sixth, 0, is name[5], which is name[name.size()]. A vector would have read past its block there, as Module 2 showed. A string has its end marker in that slot, and the standard says s[s.size()] reads it. So Bob's loop prints a 0 he did not want, with no crash and no message.

One step further is real trouble. name[name.size() + 1] is outside the text, and nothing checks it. On Compiler Explorer, at the Playground's flags, name[name.size() + 5] read 9 in one run and -15 in the next. Those are leftover bytes, so any run may give anything. So write i < s.size(), and remember that the slot at s.size() is the end marker, not a character.

size() is the same unsigned type as a vector's, so s.size() - 1 on an empty string is 18446744073709551615, exactly as Module 2, lesson 01 showed. Check empty() before you subtract.

A string is bytes: Maria's name in Bangla

Maria asks why her name has a different length when she types it in Bangla. The answer is the most honest line in this module: a string holds bytes, and size() counts bytes. A byte is 8 bits, a number from 0 to 255. Every ASCII character, the English letters, digits and punctuation, is one byte. A Bangla letter is not.

The Playground stores text in UTF-8, the encoding almost every web page and Linux system uses. UTF-8 gives every character in every language a number, its code point, written like U+09AE. Then it writes that number as one to four bytes. ASCII code points take one byte. Each Bangla code point takes three.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string en = "Maria";
    string bn = "মারিয়া";
    cout << en << ": size " << en.size() << '\n';
    cout << bn << ": size " << bn.size() << '\n';
    cout << "bytes:";
    for (unsigned char c : bn) {
        cout << ' ' << (int)c;
    }
    cout << '\n';
    return 0;
}
Maria: size 5
মারিয়া: size 21
bytes: 224 166 174 224 166 190 224 166 176 224 166 191 224 166 175 224 166 188 224 166 190

Five letters, size 5. The Bangla name looks like three letters, মা, রি and য়া, and its size is 21. It is seven code points: three letters, three vowel signs and one nukta dot, each written as three bytes. Every byte is above 127, so none of them is an ASCII character on its own. The loop uses unsigned char so the bytes print as 0 to 255; a plain char on the Playground is signed and would print them as negative numbers. Step through the name below, one code point at a time.

Maria's name in Bangla, built one code point at a time, with size() after each step. Measured on GCC 12 at the Playground's flags.

StepCode pointWhat it isIts bytessize() after
1U+09AEম, the letter ma224 166 1743
2U+09BEা, the vowel sign aa224 166 1906
3U+09B0র, the letter ra224 166 1769
4U+09BFি, the vowel sign i224 166 19112
5U+09AFয, the letter ya224 166 17515
6U+09BC়, the nukta dot that turns য into য়224 166 18818
7U+09BEা, the vowel sign aa224 166 19021

So three numbers describe one Bangla word: the bytes size() counts, the code points, and the letters a reader sees. For English text they are all the same number, which is why the problems in this module use ASCII input. For Bangla, s[0] is one byte of a letter, never the whole letter.

Joining with +, and comparing with ==

+ makes a new string from two others, and += adds to the end of an existing one. == asks whether two strings hold the same text. In C, == on two char arrays compared their addresses, and you needed strcmp. This program shows both behaviours side by side.

#include <iostream>
#include <string>
using namespace std;

int main() {
    char a[] = "tea";
    char b[] = "tea";
    cout << "C arrays equal? " << (a == b) << '\n';

    string x = "tea";
    string y = "te";
    y += 'a';
    cout << "strings equal? " << (x == y) << '\n';

    string order = x + " and " + y;
    cout << order << '\n';
    return 0;
}
C arrays equal? 0
strings equal? 1
tea and tea

The two arrays hold the same letters but live at two addresses, so a == b is false. The strings compare their characters, so x == y is true even though y was built in two steps. += took a single char, 'a', as easily as a whole string. So == means "same text" for strings, which is what you meant all along.

One rule about +: a string must be on one side of each join. x + " and " + y works because it starts from a string, and each step gives a string back. Two quoted texts are not strings; they are C arrays, so "Maria" + " " + last fails at every command line. GCC 12 says error: invalid operands of types 'const char [6]' and 'const char [2]' to binary 'operator+'. Start from a string, or write string("Maria").

Example 1: the smallest string program

Read one word, then print it, its size, and its first and last characters.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string dish;
    cin >> dish;
    cout << dish << " has " << dish.size() << " letters\n";
    cout << "first " << dish[0] << ", last " << dish[dish.size() - 1] << '\n';
    return 0;
}
biryani has 7 letters
first b, last i

That output is for the input biryani. dish.size() - 1 is safe only because the input has a word. On an empty input cin >> fails, dish stays empty, and the last index wraps round. A program for real users checks dish.empty() first.

Run in Compiler
Example 2: Amara's name card, line by line

Amara reads full names, one per line, until the input ends. For each she prints the name and its length, and at the end the longest name. The names have spaces, so she reads lines, not words.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string line, longest;
    while (getline(cin, line)) {
        cout << line << " (" << line.size() << ")\n";
        if (line.size() > longest.size()) {
            longest = line;
        }
    }
    cout << "longest: " << longest << '\n';
    return 0;
}
Maria Lopez (11)
Bob (3)
Kenji Watanabe (14)
longest: Kenji Watanabe

That output is for three lines: Maria Lopez, Bob and Kenji Watanabe. while (getline(cin, line)) stops at the end of the input, as while (cin >> x) did. longest = line; copies the text, so the next getline cannot change it.

Run in Compiler
Example 3: Zara's password check

Zara checks a password the way a sign-up form does. It must be at least 8 characters, and it must have a digit and a capital letter. She tests the empty line first, as always, so the program handles it with no special case.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string pw;
    while (getline(cin, pw)) {
        bool digit = false, capital = false;
        for (char c : pw) {
            if (c >= '0' && c <= '9') {
                digit = true;
            }
            if (c >= 'A' && c <= 'Z') {
                capital = true;
            }
        }
        bool ok = pw.size() >= 8 && digit && capital;
        cout << "[" << pw << "] " << (ok ? "strong" : "weak") << '\n';
    }
    return 0;
}
[tea4two] weak
[Tea4two!] strong
[] weak
[Biryani2026] strong

That output is for four lines, the third one empty. The range-for walks every char of the string, the way it walked a vector's elements, and an empty string gives it nothing to walk. The checks compare characters with ASCII ranges, a habit from the C track. A Bangla letter is three bytes above 127, so it counts as neither a digit nor a capital.

Run in Compiler

Where this is used

  • Protocol Buffers. Google's message format turns every string field of a .proto file into a std::string in the generated C++ code, read with an accessor such as name(). A field can hold any length, so the generated code never picks a size.
  • Chromium. The browser's URL class, GURL, keeps the whole address as a std::string and hands it out through spec(). A URL's length is only known when it is typed or clicked.
  • nlohmann/json. This widely used JSON library for C++ stores every JSON string value as a std::string by default, its string_t type. It learns each value's length only as it reads the text.

Common mistakes

1. Reading a full name with cin >>.

string name;
cin >> name;
cout << "Hello, " << name << "!\n";

No message. For the input Maria Lopez it prints Hello, Maria!. >> reads one word, and "Lopez" waits in the input for the next read. Use getline(cin, name) for anything that can hold a space. You will make this mistake because >> worked for every number you ever read.

2. Giving a string to printf's %s.

string name = "Maria";
printf("%s\n", name);

At the Playground's flags it compiles with no message, and two runs on Compiler Explorer printed two different sets of garbage characters. %s wants a C string's address, and a string is a different thing. With -Wall -Wextra on your own machine, GCC 12 says it: warning: format '%s' expects argument of type 'char*', but argument 2 has type 'std::string' {aka 'std::__cxx11::basic_string<char>'} [-Wformat=]. Print with cout, or pass name.c_str() (lesson 02). You will make it because printf is the C track's habit.

3. Writing into an index that does not exist yet.

string s;
s[0] = 'M';
cout << "[" << s << "] size " << s.size() << '\n';

No message, and one run on Compiler Explorer printed [] size 0. The write landed on the end marker of an empty string, which is undefined behaviour, and the size never changed. [] never adds a character. Use s += 'M'; or s.push_back('M');, or make room first with string s(1, ' ');. You will make it because an array of the right size already had its boxes.

Brain teaser

Kenji wants a string of three a's. Bob wants the text "31". They write these two lines.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string kenji(3, 'a');
    string bob = "3" + 1;
    cout << "[" << kenji << "] [" << bob << "]\n";
    return 0;
}

Both lines compile with no message. What does the program print, and which of them got what they wanted?

Is "3" a string? Look at what + joins in "Joining with +, and comparing with ==", and at what "3" was in C.

Exercise 1Easy

Kenji is making a word game and needs the longest word in a list. Read words until the input ends. Print the longest word and its length. If several words share the longest length, print the first of them.

Input. Words separated by spaces or line breaks, until the input ends. No count comes first.

Output. One line: the longest word, a space, and its length. On a tie, the first of the longest words.

Constraints. 1 to 100000 words, each 1 to 100 ASCII characters; the input is at most 1000000 bytes.

Sample. Input Bob packs maps, snacks and a camera for the trip gives snacks 6.

#include <iostream>
#include <string>
using namespace std;

int main()
{
    ios::sync_with_stdio(false);
    cin.tie(nullptr);

    string word;
    while (cin >> word) {
        // Keep the longest word seen so far.
        // Replace it only when this word is strictly longer.
    }

    // Print the longest word, a space, and its length.

    return 0;
}

Graded as longest-word, a free problem in this module's problem set. The hidden tests include a single word, many words of equal length, and 100000 words. They catch a program that keeps the last longest word instead of the first.

Run in Compiler
Exercise 2Medium

David prints name badges. Read one full name as a line. Print it inside a frame of stars: a line of stars, the name with * before it and * after it, and the line of stars again. The frame is exactly as wide as the middle line.

Input. One line, the name, which may contain spaces. ASCII only.

Output. Three lines: the stars, the framed name, the stars.

Constraints. The name has 1 to 100 characters.

Sample. Input Maria Lopez gives ***************, then * Maria Lopez *, then ***************. The middle line is 11 + 4 = 15 characters.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string name;
    getline(cin, name);

    // Make a string of stars as long as the framed line,
    // then print the three lines.

    return 0;
}

Not graded on its own. Hint: string(n, '*') makes n stars, and the framed line is the name's size plus four.

Run in Compiler
Exercise 3Hard

Maria wants a program that counts code points, not bytes, so that "মারিয়া" gives 7 and "Maria" gives 5. Read one line of UTF-8 text and print the number of code points in it.

Input. One line of UTF-8 text, English, Bangla or both.

Output. One number, the count of code points.

Constraints. The line has 1 to 100000 bytes.

Sample. Input Maria মারিয়া gives 13: five letters, one space and seven code points.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string line;
    getline(cin, line);

    // Count the bytes that start a code point.
    // Hint: look at the bytes of the Bangla name in this lesson.

    return 0;
}

Not graded on its own. The idea the lesson only pointed at: in UTF-8, the second and third bytes of a Bangla code point are always between 128 and 191. A byte that starts a code point never is. Count the other bytes, using unsigned char.

Run in Compiler

Common doubts

  • Should I write size() or length()?

    They return the same number, and both cost nothing. length() sounds right for text, but every container in the STL has size(), so code that uses size() reads the same across them. This track uses size().

  • Do I still need char arrays?

    Rarely. A C library function may want a const char*, and lesson 02 shows c_str(), which gives it one. Lesson 04 lists the few places where something other than a string is the better choice.

  • Is a string just a vector<char>?

    Inside it looks much the same: characters in one block, a size and a capacity. On top it adds what text needs: the end marker, + to join, getline to read, and find and substr to search and cut, which lesson 02 lists.

  • My Bangla name printed a different size on another computer. Why?

    The same letter can be stored in more than one way. য় can be one code point, U+09DF, or two, য followed by the nukta dot, as in this lesson. With U+09DF the name is 18 bytes, not 21. The screen shows the same letters either way, and lesson 05 (Pro) measures both.

Key takeaways

  • A string is a row of characters that grows by itself and keeps its size, so you never pick a length or call strlen.
  • cin >> s reads one word; getline(cin, s) reads a whole line, spaces included.
  • Indexes run from 0 to s.size() - 1; s[s.size()] is the end marker '\0', and past it nothing is checked.
  • + and += join, = copies the text, and == compares the text, not the address.
  • size() counts bytes: an ASCII character is one, a Bangla code point is three, so "মারিয়া" has size 21.
  • Go deeper: Under the Hood, the small string buffer and the cost of + (Pro).

Next, Bob builds a long line with s = s + c and waits. Lesson 02 puts a cost on every string operation and shows why.

End of lesson 1

Mark it done, and your progress moves with you.

Next: Every string Operation, One by One, With Its Cost