Learn C++ STL

Lesson 2 of 9 · string: Text That Knows Its Own Length

Module 3 · string: Text That Knows Its Own Length

Every string Operation, One by One, With Its Cost

FreeReading

In this lesson

  • Use every string operation a working programmer calls, from size to ostringstream, with its exact syntax.
  • State the cost of each one and explain why s += c is cheap while s = s + c in a loop is slow.
  • Handle the three ways text goes wrong: a position past the end, a search that finds nothing, and a number that is not a number.

Bob needs one line of 100,000 letters for a test file. He writes s = s + c; in a loop, and the program sits there for nearly half a second, sometimes over a whole second. Amara writes s += c; instead, and hers finishes in under a millisecond. The two lines look the same and print the same string. The difference is one cost line of this lesson, and the chart at the end shows it measured.

How to read a cost line

Lesson 01 showed what a string is: a row of characters that knows its own length. Here every operation ends with a one-row table: the call, its cost, and why. The cost is written in big-O notation, as in Module 2. It says how the work grows as the text grows, not how many nanoseconds it takes.

CostRead it asExample
O(1)the same small amount of work at any lengths.size() on 5 characters or 5 million
amortised O(1)constant on average over many calls; now and then one call copies everythings += c, which sometimes grows the block
O(n)work that grows with the n characters it touchess.insert(0, "x") moves every character
O(n x m)for each of n places, up to m comparisonssearching for an m-letter word in n characters, at worst

In this lesson n is the length of the string you call the operation on, and m is the length of the other piece. So a cost line tells you what happens when the text is a whole book instead of a name.

size, length and empty

How long is it?

s.size()      the number of characters, as a size_t
s.length()    exactly the same number as size()
s.empty()     true when size() is 0

A size_t is the unsigned whole-number type the library uses for sizes and positions. It cannot hold a negative number.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string city = "Chattogram";
    string none;
    cout << city.size() << ' ' << city.length() << ' ' << city.empty() << '\n';
    cout << none.size() << ' ' << none.empty() << '\n';
    cout << none.size() - 1 << '\n';
    return 0;
}
10 10 0
0 1
18446744073709551615
CallCostBecause
s.size(), s.length(), s.empty()O(1)the string stores its length; nothing is counted, unlike C's strlen

Watch: the last line. An empty string's size() - 1 is not -1; an unsigned number wraps around to the largest size_t. So check !s.empty() before you subtract from size().

operator[] and at

One character by its index

s[i]       the character at index i, no check
s.at(i)    the same character, checked: throws std::out_of_range if i >= s.size()

Both give a char you can read or change. Lesson 01 showed what s[i] does past the end. at checks the index first. When the index is too big, it throws an exception: it stops the normal flow and reports an error object.

#include <iostream>
#include <stdexcept>
#include <string>
using namespace std;

int main() {
    string name = "Maria";
    name[0] = 'm';
    name.at(4) = 'A';
    cout << name << ' ' << name[1] << ' ' << name.at(2) << '\n';
    try {
        cout << name.at(5) << '\n';
    } catch (const out_of_range& e) {
        cout << "caught: " << e.what() << '\n';
    }
    return 0;
}
mariA a r
caught: basic_string::at: __n (which is 5) >= this->size() (which is 5)

try { ... } catch (...) { ... } runs the first block. If something inside it throws an out_of_range, the program jumps to the catch block instead of stopping. e.what() is the message the library wrote. Without the try, the same call ends the program. On Compiler Explorer the run stopped with exit code 134, and the error stream said this:

cout << name.at(5) << '\n';
terminate called after throwing an instance of 'std::out_of_range'
  what():  basic_string::at: __n (which is 5) >= this->size() (which is 5)
CallCostBecause
s[i], s.at(i)O(1)the character's address is the start plus i bytes; at adds one comparison

Watch: name[5] reads the hidden '\0' that lesson 01 drew, while name.at(5) throws. So at allows indexes 0 to size() - 1 only, and says so loudly.

front and back

The two ends

s.front()    the first character, the same as s[0]
s.back()     the last character, the same as s[s.size() - 1]
#include <iostream>
#include <string>
using namespace std;

int main() {
    string word = "radar";
    cout << word.front() << word.back() << '\n';
    word.front() = 'R';
    cout << word << '\n';
    return 0;
}
rr
Radar
CallCostBecause
s.front(), s.back()O(1)one read at a known address

Watch: Zara's first test is the empty string. There, both calls are undefined behaviour. In one run on Compiler Explorer, each quietly gave 0, and nothing stopped the program. So check !s.empty() first, because no error will tell you.

+=, append, push_back and pop_back

Changing the end

s += t;             add a string, a literal or one char at the end
s.append(t);        the same as s += t
s.append(n, c);     add n copies of the char c
s.push_back(c);     add one char at the end
s.pop_back();       remove the last char, return nothing
#include <iostream>
#include <string>
using namespace std;

int main() {
    string line = "Hello";
    line += ',';
    line += " Zara";
    line.append(3, '!');
    line.push_back('?');
    cout << line << '\n';
    line.pop_back();
    cout << line << " (" << line.size() << ")\n";
    return 0;
}
Hello, Zara!!!?
Hello, Zara!!! (14)
CallCostBecause
s += c, s.push_back(c)amortised O(1)the string keeps spare room at the end; when it runs out, it doubles the block and copies once
s += t, s.append(t)amortised O(m)the m new characters are copied in; the old ones stay where they are
s.pop_back()O(1)the length drops by one

Watch: pop_back() on an empty string is undefined. In one run on Compiler Explorer, size() afterwards printed 72057594037927935, and the run ended normally. So pop_back needs the same empty() check as back.

+, which builds a new string

Joining into a new string

a + b        a new string: the characters of a, then those of b

Lesson 01 joined two strings with +. One side must be a string; the other may be a string, a literal or a char. Neither side changes.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string first = "Amara";
    string last = "Okafor";
    string full = first + ' ' + last;
    cout << full << " (" << full.size() << ")\n";
    cout << first << '\n';
    return 0;
}
Amara Okafor (12)
Amara
CallCostBecause
a + bO(n + m)a new block is made and both sides are copied into it

Watch: this is Bob's slow line. s = s + c copies all of s into a new string just to add one character. In a loop of n steps, that is about n2 / 2 copied characters. s += c adds the character in place. So use + to build a value once, and += to grow a string.

==, !=, < and compare

Comparing two strings

a == b, a != b              same characters in the same order?
a < b, a <= b, a > b, a >= b  dictionary order, byte by byte
a.compare(b)                negative, 0 or positive: a before, equal to, or after b

Dictionary order compares the first characters, then the second, and so on, until two differ. The smaller byte value wins. If one string runs out first, the shorter one comes first.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string a = "Zara";
    string b = "apple";
    cout << (a == "Zara") << ' ' << (a != b) << ' ' << (a < b) << '\n';
    cout << (string("app") < "apple") << '\n';
    int r = a.compare(b);
    if (r < 0) {
        cout << a << " comes before " << b << '\n';
    } else if (r > 0) {
        cout << a << " comes after " << b << '\n';
    } else {
        cout << "equal\n";
    }
    return 0;
}
1 1 1
1
Zara comes before apple

Maria asks why "Zara" comes before "apple". The answer is the ASCII table: 'Z' is 90 and 'a' is 97. Every capital letter is smaller than every small letter.

CallCostBecause
==, <, compareO(n) at mostthe bytes are compared from the start, and the work stops at the first difference

Watch: compare promises only the sign of its answer. Test r < 0, never r == -1. So read compare by its sign, and remember that string order is byte order, not the order a dictionary prints.

substr

A piece of the string, as a new string

s.substr(pos, len)   len characters starting at index pos
s.substr(pos)        from pos to the end
#include <iostream>
#include <stdexcept>
#include <string>
using namespace std;

int main() {
    string date = "2026-10-07";
    cout << date.substr(0, 4) << '|' << date.substr(5, 2) << '|' << date.substr(8) << '\n';
    cout << date.substr(8, 100) << '\n';
    cout << '[' << date.substr(10) << "]\n";
    try {
        cout << date.substr(11) << '\n';
    } catch (const out_of_range& e) {
        cout << "caught: " << e.what() << '\n';
    }
    return 0;
}
2026|10|07
07
[]
caught: basic_string::substr: __pos (which is 11) > this->size() (which is 10)

The second number is a length, not an end index. A length that runs past the end is cut short, as substr(8, 100) shows. A start equal to size() gives an empty string. A start past size() throws out_of_range.

CallCostBecause
s.substr(pos, len)O(len)the piece is copied into a new string

Watch: substr makes a copy. Changing the piece never changes s. So a substr inside a loop over a long text costs a copy every time.

In C++20

starts_with and ends_with test a prefix (the start) or a suffix (the end) in one call, without the copy that s.substr(0, 8) == "https://" makes. Both exist on string and on string_view, the C++17 type that looks at characters without owning them (lesson 04 teaches it). On the Playground's C++17 the same line stops with error: 'std::string' {aka 'class std::__cxx11::basic_string<char>'} has no member named 'starts_with'. The Run button below opens the Playground at C++20.

#include <iostream>
#include <string>
#include <string_view>
using namespace std;

int main() {
    string url = "https://progsity.io/learn";
    cout << url.starts_with("https://") << ' ' << url.ends_with(".pdf") << '\n';
    string_view file = "report.pdf";
    cout << file.starts_with('r') << ' ' << file.ends_with(".pdf") << '\n';
    return 0;
}
1 0
1 1

Each call compares only as many characters as the prefix or suffix has, so its cost is O(m).

Run in Compiler

find, rfind and string::npos

Searching

s.find(x)         index of the first x (a char or a string), or string::npos
s.find(x, pos)    the same, starting the search at index pos
s.rfind(x)        index of the last x, or string::npos
string::npos      "no position": the largest value a size_t can hold
#include <iostream>
#include <string>
using namespace std;

int main() {
    string path = "docs/stl/string.html";
    cout << path.find('/') << ' ' << path.rfind('/') << '\n';
    cout << path.find("stl") << ' ' << path.find('/', 5) << '\n';
    size_t hash = path.find('#');
    if (hash == string::npos) {
        cout << "no #\n";
    }
    cout << string::npos << '\n';
    return 0;
}
4 8
5 8
no #
18446744073709551615
CallCostBecause
s.find(c), s.rfind(c)O(n)each character is checked once, from one end
s.find(t), s.rfind(t)up to O(n x m) on GCC's libraryat each candidate place, up to m characters are compared

Watch: find never returns -1. When it finds nothing, it returns string::npos, 18446744073709551615 on the Playground. Keep the result in a size_t and compare it with string::npos, as the program does. So "not found" is a huge number, not a negative one.

find_first_of and find_last_of

Searching for any character of a set

s.find_first_of(set)       index of the first char that is in set, or npos
s.find_last_of(set)        index of the last char that is in set, or npos
s.find_first_not_of(set)   index of the first char that is NOT in set, or npos
s.find_last_not_of(set)    index of the last char that is NOT in set, or npos
#include <iostream>
#include <string>
using namespace std;

int main() {
    string line = "price: 450 taka";
    size_t first = line.find_first_of("0123456789");
    size_t last = line.find_last_of("0123456789");
    cout << first << ' ' << last << ' ' << line.substr(first, last - first + 1) << '\n';

    string padded = "   Zara  ";
    size_t b = padded.find_first_not_of(' ');
    size_t e = padded.find_last_not_of(' ');
    cout << '[' << padded.substr(b, e - b + 1) << "]\n";
    return 0;
}
7 9 450
[Zara]

The second half trims the string: it cuts the spaces off both ends. The _not_of pair is how you find where the real text starts and stops.

CallCostBecause
find_first_of(set) and its three siblingsup to O(n x k) for a set of k characterseach character of the string is checked against the set

Watch: the argument is a set of characters, not a word. Bob's s.find_first_of("cat") stops at the first c, a or t anywhere. To find the word, use find.

insert and erase

Adding and removing in the middle

s.insert(pos, t)       put the string t before index pos
s.insert(pos, n, c)    put n copies of the char c before index pos
s.erase(pos, len)      remove len characters starting at pos
s.erase(pos)           remove everything from pos to the end
#include <iostream>
#include <string>
using namespace std;

int main() {
    string s = "Dhaka 2026";
    s.insert(5, " city");
    cout << s << '\n';
    s.insert(0, 1, '[');
    s.append(1, ']');
    cout << s << '\n';
    s.erase(6, 5);
    cout << s << '\n';
    s.erase(6);
    cout << s << '\n';
    return 0;
}
Dhaka city 2026
[Dhaka city 2026]
[Dhaka 2026]
[Dhaka
CallCostBecause
s.insert(pos, t)O(n - pos + m), so O(n) at the frontevery character after pos shifts right to make room
s.erase(pos, len)O(n - pos)every character after the gap shifts left to close it

Watch: s.erase(6) with one number removes everything from index 6 on, not one character. To remove one character, write s.erase(6, 1). So with erase, the second number is a length, as with substr.

replace and clear

Overwriting a piece, and emptying

s.replace(pos, len, t)   remove len characters at pos and put t there
s.clear()                remove every character; size() becomes 0
#include <iostream>
#include <string>
using namespace std;

int main() {
    string msg = "Meet at 5 pm at the gate";
    msg.replace(8, 4, "6:30 pm");
    cout << msg << '\n';
    size_t at = msg.find("gate");
    msg.replace(at, 4, "library");
    cout << msg << '\n';
    msg.clear();
    cout << '[' << msg << "] " << msg.size() << ' ' << msg.empty() << '\n';
    return 0;
}
Meet at 6:30 pm at the gate
Meet at 6:30 pm at the library
[] 0 1

The new piece may be longer or shorter than the old one. "5 pm" had 4 characters and "6:30 pm" has 7, so the rest of the line moved right by 3.

CallCostBecause
s.replace(pos, len, t)O(n - pos + m)the tail moves when the lengths differ, and m characters are written
s.clear()O(1)a char needs no clean-up, so only the length changes; the memory stays

Watch: replace changes one place, the one you name by index. It does not search. Replacing every copy of a word needs a loop with find, and Exercise 3 asks for exactly that.

c_str and data

The characters as a C string

s.c_str()    a const char* to the characters, ending in '\0'
s.data()     the same pointer as c_str()

Some functions come from C and want a const char*, a pointer to characters that end in '\0'. printf with %s and strlen are two of them.

#include <cstdio>
#include <cstring>
#include <string>
using namespace std;

int main() {
    string name = "Kenji";
    const char* p = name.c_str();
    printf("%s has %d letters\n", p, (int)strlen(p));
    printf("%s again\n", name.data());
    return 0;
}
Kenji has 5 letters
Kenji again
CallCostBecause
s.c_str(), s.data()O(1)the string already keeps its characters with a '\0' after them; it hands out the address

Watch: the pointer is valid only until the string changes. After a += that grows the block, the characters may live somewhere new. So call c_str() right where you need it, and never keep the pointer.

stoi, stol, stoll and stod

Text to a number

stoi(s)          an int      stol(s)    a long
stoll(s)         a long long stod(s)    a double
stoi(s, &used)   the same, and used becomes the number of characters read

Each one skips leading spaces, reads as many characters as form a number, and stops at the first one that does not.

#include <iostream>
#include <stdexcept>
#include <string>
using namespace std;

int main() {
    cout << stoi("42") + 1 << ' ' << stoi("  -7") << '\n';
    cout << stoi("12abc") << '\n';
    size_t used = 0;
    int v = stoi("12abc", &used);
    cout << v << " used " << used << '\n';
    cout << stoll("9000000000") << ' ' << stod("3.75") * 2 << '\n';
    try {
        cout << stoi("abc") << '\n';
    } catch (const invalid_argument& e) {
        cout << "invalid_argument: " << e.what() << '\n';
    }
    try {
        cout << stoi("3000000000") << '\n';
    } catch (const out_of_range& e) {
        cout << "out_of_range: " << e.what() << '\n';
    }
    return 0;
}
43 -7
12
12 used 2
9000000000 7.5
invalid_argument: stoi
out_of_range: stoi

stoi("12abc") gives 12 with no complaint, and used says only 2 characters were read. Text with no number at the start throws invalid_argument. A number too big for the type throws out_of_range. In both cases what() is only the function's name.

CallCostBecause
stoi, stol, stoll, stodO(d) for d characters readeach digit is read once

Watch: on 64-bit Linux, where the Playground runs, a long is 8 bytes, the same as a long long. On 64-bit Windows it is 4. Use stoll for big whole numbers, and it means the same everywhere. So pick the function by the size of the number, and catch the two exceptions when the text comes from a user.

to_string

A number to text

to_string(x)   x as a string; x may be int, long long, double and the rest
#include <iostream>
#include <string>
using namespace std;

int main() {
    int marks = 87;
    string line = "Zara: " + to_string(marks) + "/100";
    cout << line << " (" << line.size() << ")\n";
    cout << to_string(-5) << ' ' << to_string(2.5) << '\n';
    return 0;
}
Zara: 87/100 (12)
-5 2.500000
CallCostBecause
to_string(x)O(d) for d digitseach digit is written once into a new string

Watch: a double always comes out with six decimals, so 2.5 becomes 2.500000. For a number printed the way cout prints it, use an ostringstream, two sections below.

getline with a delimiter

Read up to a chosen character

getline(in, s)         read up to '\n', which is dropped
getline(in, s, ',')    read up to ',', which is dropped; '\n' is now an ordinary char

Lesson 01 read a whole line with getline(cin, s). A third argument, a char, changes where it stops. That character is the delimiter, the mark between one piece and the next. Here in is any input stream: cin, or the string streams below.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string name;
    while (getline(cin, name, ',')) {
        cout << '[' << name << "]\n";
    }
    return 0;
}
[Alice]
[Bob]
[Maria
]

That output is for the input Alice,Bob,Maria on one line. Look at the last name: the closing bracket is on the next line.

CallCostBecause
getline(in, s, c)O(k) for the k characters readeach one is read once and added with amortised O(1)

Watch: with a delimiter, the end of the line is just another character, so the last field kept its '\n'. The clean way is two steps: read the line with getline(cin, line), then split the line. So the next section is the tool for the second step.

istringstream: values out of a line

A stream that reads from a string

#include <sstream>
istringstream in(line);      a stream over a copy of line
in >> x                      read the next value, exactly as cin >> x
while (in >> x)              stops at the end, or at text that is not an x
getline(in, field, ',')      split the line on a char

Module 1's lesson on fast input pointed at istringstream: any >> loop that works on cin works on a string. Here is the whole tool. The stream keeps its own position in the line, so each read starts where the last one stopped.

#include <iostream>
#include <sstream>
#include <string>
using namespace std;

int main() {
    istringstream row("Amara 85 90 77");
    string name;
    int a, b, c;
    row >> name >> a >> b >> c;
    cout << name << " total " << a + b + c << '\n';

    istringstream nums("4 8 x 15");
    int x, sum = 0;
    while (nums >> x) {
        sum += x;
    }
    cout << "sum " << sum << '\n';

    istringstream cities("Dhaka,Sylhet,Khulna");
    string city;
    while (getline(cities, city, ',')) {
        cout << '[' << city << "]\n";
    }
    return 0;
}
Amara total 252
sum 12
[Dhaka]
[Sylhet]
[Khulna]

The sum is 12, not 27. When >> met x, it could not read an int, so the stream failed: it marked itself as broken, and every later read gives false. The loop stopped there, and 15 was never read.

CallCostBecause
istringstream in(line)O(n)the stream keeps its own copy of the line
in >> x, getline(in, field, c)O(k) for the k characters readreading moves forward and never goes back

Watch: a failed stream stays failed. Make a new istringstream for each line rather than reusing one. So a line split is one fresh stream, read until it runs dry.

ostringstream: text built with <<

A stream that writes into a string

#include <sstream>
ostringstream out;     an empty stream
out << x << " text";   write values, exactly as cout << x
out.str()              the whole text so far, as a string
#include <iostream>
#include <sstream>
#include <string>
using namespace std;

int main() {
    ostringstream out;
    out << "Bob ran " << 5 << " km in " << 27.5 << " min";
    string report = out.str();
    cout << report << '\n';
    cout << report.size() << " characters\n";
    return 0;
}
Bob ran 5 km in 27.5 min
24 characters

The number 27.5 came out as 27.5, the way cout prints it, not as to_string's 27.500000. Anything you can print with cout, you can write into an ostringstream, and keep as a string.

CallCostBecause
out << xamortised O(k) for the k characters writtenthe stream grows its buffer the way += does
out.str()O(n)it returns a copy of everything written so far

Watch: str() copies. Call it once at the end, not after every <<. So the pattern is: write everything into one stream, then take the string once.

reserve and capacityIntermediate

Room before you need it

s.capacity()    how many characters fit before the block must grow
s.reserve(n)    make the capacity at least n; the size does not change
#include <iostream>
#include <string>
using namespace std;

int main() {
    string s;
    cout << s.size() << ' ' << s.capacity() << '\n';
    s.reserve(100);
    cout << s.size() << ' ' << s.capacity() << '\n';
    for (int i = 0; i < 60; i++) {
        s += 'x';
    }
    cout << s.size() << ' ' << s.capacity() << '\n';
    return 0;
}
0 15
0 100
60 100

An empty string already had room for 15 characters on GCC 12. Why 15, and where those characters live, is lesson 05's story. After reserve(100), 60 appends never had to grow the block.

CallCostBecause
s.capacity()O(1)the string stores it
s.reserve(n)O(n) when it grows the block, O(1) when it does nota new block means copying every character across

Watch: reserve makes room, not characters; s[0] after s.reserve(100) is still past the end. So reserve helps only when you know the final length in advance.

Every cost in one table, and Bob's measurement

OperationCost
size, length, empty, [], at, front, back, pop_back, c_str, data, capacity, clearO(1)
+= one char, push_backamortised O(1)
+= or append of m charactersamortised O(m)
a + bO(n + m), a new string
==, <, compareO(n) at most
substr(pos, len)O(len), a copy
find, rfind of a char; of an m-char stringO(n); up to O(n x m)
find_first_of and its siblings, a set of k charsup to O(n x k)
insert, erase, replace at posO(n - pos) plus the characters written
stoi, stoll, stod, to_stringO(d) for d digits
getline, in >> x, out << xO(k) for the k characters moved
reserveO(n) if it grows the block

Here is Bob's line measured. The program builds a 100,000-character string and times only the loop with <chrono>'s steady_clock. We ran it on Compiler Explorer, GCC 12 at the Playground's -O2 -std=c++17, seven times. Two more versions changed only the line in the loop, to s += c; and to s.insert(0, 1, c);, and each ran seven times too.

#include <chrono>
#include <iostream>
#include <string>
using namespace std;

int main() {
    const int n = 100000;
    auto start = chrono::steady_clock::now();
    string s;
    for (int i = 0; i < n; i++) {
        char c = 'a' + i % 26;
        s = s + c;
    }
    auto stop = chrono::steady_clock::now();
    chrono::duration<double, milli> took = stop - start;
    cout << s.size() << " characters in " << took.count() << " ms\n";
    return 0;
}

Each run prints one line, the size and then the time in milliseconds, and the time moved from run to run. The chart shows the middle run of the seven as a bar, and the lowest and highest as the line across it.

Building a 100,000-character string: s += c against s.insert(0, 1, c) and s = s + c Time to build 100,000 characters, GCC 12, -O2 -std=c++17, seven runs each Bar: the middle run. Line: the fastest to the slowest run. Each grid step is ten times longer. s += c 0.37 ms (0.29 to 0.72) s.insert(0, 1, c) 73 ms (70 to 96) s = s + c 416 ms (402 to 1142) 0.1 ms 1 ms 10 ms 100 ms 1 s 10 s Same output, same O(n) per step for the last two; s = s + c also makes and frees a new string every step.

The += loop never took a whole millisecond. The + loop was at least 500 times slower in every run. Inserting at the front is O(n) per call too, yet it beat s = s + c, because it shifts the characters inside the same block. The + loop also makes a new block and frees the old one, 100,000 times. So both cost lines say O(n2) for the whole loop, and the measurement shows the extra price of a new string each time.

Example 1: David splits a web address

David's tool prints the parts of a web address. find gives each boundary, and substr cuts between them. The second address has no ?, which is Zara's case.

#include <iostream>
#include <string>
using namespace std;

void show_parts(const string& url) {
    size_t scheme_end = url.find("://");
    size_t host_start = scheme_end + 3;
    size_t path_start = url.find('/', host_start);
    size_t query_start = url.find('?', path_start);

    cout << "scheme " << url.substr(0, scheme_end) << '\n';
    cout << "host   " << url.substr(host_start, path_start - host_start) << '\n';
    cout << "path   " << url.substr(path_start, query_start - path_start) << '\n';
    if (query_start == string::npos) {
        cout << "query  (none)\n";
    } else {
        cout << "query  " << url.substr(query_start + 1) << '\n';
    }
}

int main() {
    show_parts("https://progsity.io/learn/stl?lang=bn");
    show_parts("http://example.com/about");
    return 0;
}
scheme https
host   progsity.io
path   /learn/stl
query  lang=bn
scheme http
host   example.com
path   /about
query  (none)

For the second address, query_start is npos, so the path's length is enormous. substr cuts a too-long length at the end, so the path is simply the rest. That rule did the work of an extra if.

Run in Compiler
Example 2: Alice adds up her study sessions

Alice logs each study session as a start and an end time, one session per line. A small function turns "07:45" into minutes after midnight with find, substr and stoi. An istringstream splits each line, and an ostringstream collects the report, printed once.

#include <iostream>
#include <sstream>
#include <string>
using namespace std;

int to_minutes(const string& hhmm) {
    size_t colon = hhmm.find(':');
    int h = stoi(hhmm.substr(0, colon));
    int m = stoi(hhmm.substr(colon + 1));
    return h * 60 + m;
}

int main() {
    ostringstream out;
    string line;
    int total = 0;
    while (getline(cin, line)) {
        istringstream in(line);
        string start, stop;
        in >> start >> stop;
        int minutes = to_minutes(stop) - to_minutes(start);
        total += minutes;
        out << start << " to " << stop << ": " << minutes << " min\n";
    }
    out << "total " << total / 60 << " h " << total % 60 << " min\n";
    cout << out.str();
    return 0;
}
07:45 to 09:10: 85 min
13:05 to 14:00: 55 min
20:30 to 22:15: 105 min
total 4 h 5 min

That output is for three input lines: 07:45 09:10, 13:05 14:00 and 20:30 22:15. stoi("07") is 7, because a leading zero is just a digit. The report was built in memory and printed with one cout.

Run in Compiler
Example 3: Maria finds file extensions

Maria wants the extension of each file name: the part after the last dot. rfind finds the last dot. Zara adds three edge cases first: no dot at all, a name that starts with a dot, and a name that ends with one.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string file;
    while (cin >> file) {
        size_t dot = file.rfind('.');
        if (dot == string::npos || dot == 0 || dot + 1 == file.size()) {
            cout << file << ": no extension\n";
        } else {
            cout << file << ": " << file.substr(dot + 1) << '\n';
        }
    }
    return 0;
}
report.final.pdf: pdf
notes: no extension
photo.JPG: JPG
.bashrc: no extension
archive.: no extension

That output is for the input report.final.pdf notes photo.JPG .bashrc archive. on one line. rfind skipped the first dot of report.final.pdf, which find would have stopped at. The test dot == string::npos comes first, so dot + 1 is never computed from npos.

Run in Compiler

Where this is used

  • nginx. The web server matches each request's path against its location prefixes and keeps the longest one that matches. That test is a starts_with on the path.
  • The Linux /proc files. /proc/meminfo is plain text, one line per value, such as MemTotal: followed by a number and kB. The free command reads it, and a C++ tool reads it with getline and an istringstream.
  • CSV files, RFC 4180. The format puts a field that contains a comma inside double quotes. So getline(in, field, ',') is correct only for files you know have no quoted fields, which is why real CSV readers parse the quotes too.

Common mistakes

1. Reading substr's second number as an end index.

string date = "2026-10-07";
string month = date.substr(5, 7);

No message at any command line. month is "10-07", seven characters from index 5, not "10". Write date.substr(5, 2), or date.substr(start, end - start) when you know both ends. You will make this mistake because Python's slices and many other languages take an end.

2. Calling stoi on text that may not be a number.

string age = "twenty";
int years = stoi(age);

It compiles with no message. On Compiler Explorer the run stopped with exit code 134, and the error stream said terminate called after throwing an instance of 'std::invalid_argument', then what(): stoi. Wrap the call in try and catch (const invalid_argument& e) when the text comes from a person. You will skip it because every sample input is a clean number.

3. Comparing a string with a char.

string s = "Maria";
if (s == 'M') {
    cout << "starts with M\n";
}

An error, with a page of notes after it. The first line is the one to read: error: no match for 'operator==' (operand types are 'std::string' {aka 'std::__cxx11::basic_string<char>'} and 'char'). A string is never equal to one character. Write s[0] == 'M' (after checking !s.empty()) or s == "M". You will write it because 'M' and "M" look almost the same.

4. Giving getline a string as the delimiter.

while (getline(in, city, ",")) {

An error: error: no matching function for call to 'getline(std::istringstream&, std::string&, const char [2])'. The delimiter is one char, so write ',' in single quotes. You will reach for double quotes because the rest of the line is about strings.

Brain teaser

Bob checks that a name has no 'x' in two ways.

#include <iostream>
#include <string>
using namespace std;

int main() {
    string s = "Maria";
    if (s.find('x') == -1) {
        cout << "test 1: not found\n";
    }
    unsigned int pos = s.find('x');
    if (pos == string::npos) {
        cout << "test 2: not found\n";
    }
    cout << "done\n";
    return 0;
}

On Compiler Explorer's GCC 12, at the Playground's flags, it prints test 1: not found and done, but never test 2. Explain why test 1 works although find never returns -1. Then explain why test 2 can never pass on a 64-bit machine.

Print sizeof(size_t), sizeof(unsigned int), string::npos and pos. Then ask what happens to -1 when it is compared with a size_t.

Exercise 1Easy

David's mailing list needs each address split at its @. Print the part before it and the part after it.

Input. n, then n email addresses, one per line.

Output. For each address, one line: the part before the @, a space, and the part after it.

Constraints. 1 <= n <= 20000; each address has at most 100 ASCII characters, no spaces, exactly one @, and at least one character on each side of it.

Sample. Input 3, maria@school.example, bob.rahman@mail.example.com and z@x.io gives maria school.example, bob.rahman mail.example.com and z x.io on three lines.

#include <iostream>
#include <string>
using namespace std;

int main()
{
    ios::sync_with_stdio(false);
    cin.tie(nullptr);

    int n = 0;
    cin >> n;
    for (int i = 0; i < n; i++) {
        string email;
        cin >> email;

        // Find the '@'. Print the part before it, a space,
        // and the part after it, on one line.
    }

    return 0;
}

Graded as email-domain, a free problem in this module's problem set.

Run in Compiler
Exercise 2Medium

Kenji's bank report prints big numbers, and nobody can read 1234567 at a glance. Print each number with a comma every three digits, counted from the right.

Input. n, then n integers, one per line.

Output. Each number on its own line, with its commas.

Constraints. 1 <= n <= 20000, and each integer is between -1018 and 1018.

Sample. Input 5, 0, 999, 1000, -1234567 and 1000000000000000000 gives 0, 999, 1,000, -1,234,567 and 1,000,000,000,000,000,000 on five lines.

#include <iostream>
#include <string>
using namespace std;

int main()
{
    ios::sync_with_stdio(false);
    cin.tie(nullptr);

    int n = 0;
    cin >> n;
    for (int i = 0; i < n; i++) {
        long long x = 0;
        cin >> x;

        // Turn x into text with to_string. Print it with a comma
        // between every group of three digits, counted from the right.
        // A minus sign stays in front, with no comma after it.
    }

    return 0;
}

Graded as thousands-separator, a Pro problem in this module's problem set.

Run in Compiler
Exercise 3Medium

Amara edits a line of text: every occurrence of one word must become another. Scan from left to right, and never search inside a piece you have just put in. When two matches overlap, the leftmost one wins, so replacing aa with b in aaa gives ba and a count of 1.

Input. Line 1: the text, which may contain spaces. Line 2: from. Line 3: to.

Output. The new text on line 1, and the number of replacements on line 2.

Constraints. The text has 1 to 100000 ASCII characters; from and to each have 1 to 10 characters and no space.

Sample. Input I like cats. My cat likes catnip., cat and dog gives I like dogs. My dog likes dognip. and 3 on two lines.

#include <iostream>
#include <string>
using namespace std;

int main()
{
    ios::sync_with_stdio(false);
    cin.tie(nullptr);

    string text, from, to;
    getline(cin, text);
    getline(cin, from);
    getline(cin, to);

    // Replace every occurrence of from in text, left to right.
    // Never search inside a piece you have just put in.
    // Print the new text, then the number of replacements, on two lines.

    return 0;
}

Graded as replace-word, a Pro problem in this module's problem set.

Run in Compiler
Exercise 4Hard

Zara tests a sign-up form whose age field accepts any text. For each entry, decide whether the whole entry is a whole number that fits in an int. stoi alone accepts 12abc, so it is not enough.

Input. A line with n, then n entries, one per line. Each entry has no spaces.

Output. For each entry, one line. A valid entry is an optional + or - followed by digits only, with a value that fits in an int. Print a valid entry's value as an int, and print invalid for any other entry.

Constraints. 1 <= n <= 1000. Each entry has 1 to 30 printable ASCII characters.

Sample. Input 5, then 42, -17, 12abc, 99999999999 and 007 gives 42, -17, invalid, invalid and 7.

#include <iostream>
#include <stdexcept>
#include <string>
using namespace std;

int main()
{
    ios::sync_with_stdio(false);
    cin.tie(nullptr);

    int n = 0;
    cin >> n;
    for (int i = 0; i < n; i++) {
        string entry;
        cin >> entry;

        // Print the entry's int value if the whole entry is a number
        // that fits in an int; otherwise print "invalid".
    }

    return 0;
}

Not graded on its own. The idea the lesson only pointed at: stoi can tell you how many characters it used, and a catch can turn its two exceptions into an answer.

Run in Compiler

Common doubts

  • Why is stoi better than C's atoi?

    atoi("abc") gives 0 on the Playground, so you cannot tell bad text from a real 0. stoi throws instead, and its second argument says how much it read. Both read "12abc" as 12.

  • To split a line, should I use find and substr, or a string stream?

    For values separated by spaces, istringstream with >> is the shortest. For one delimiter, getline(in, field, ',') is clean. Use find and substr when the pieces have different separators, as in David's web address. Lesson 03 writes both splits side by side.

  • How do I compare two strings ignoring capitals, so that "apple" comes before "Zara"?

    Compare lower-case copies. Lesson 03 lowers each character with tolower from <cctype>. The library's own < always compares bytes.

Key takeaways

  • size, [], at, back and c_str are O(1); at checks and throws std::out_of_range, and front, back, pop_back are undefined on an empty string.
  • s += c is amortised O(1); s = s + c copies the whole string, measured at 402 to 1142 ms for 100,000 characters.
  • substr(pos, len) takes a length, cuts a too-long one, and throws when pos > size(); erase and replace take a length too.
  • find returns string::npos for "not found"; keep it in a size_t and compare with npos.
  • stoi stops at the first non-digit and throws invalid_argument or out_of_range; istringstream reads values out of a line and ostringstream builds one.
  • Go deeper: Under the Hood, the small string buffer and the cost of + (Pro).

Next, lesson 03 builds full programs with these operations, from counting one word to a complete text report.

End of lesson 2

Mark it done, and your progress moves with you.

Next: Full Programs: From One Word to a Text Report