一:string的概念与使用
1.string到底是什么?
在 C 语言中,字符串本质上是一个以'\0'结尾的字符序列,比如:
char str[] = "hello";它实际上类似:
h e l l o \0C 语言虽然提供了strlen、strcpy、strcat等函数,但字符串的数据和操作字符串的函数是分离的,而且底层空间经常需要程序员自己管理,因此容易出现越界、空间不足等问题
C++ 中可以直接:
#include <string> using namespace std; string s = "hello";你可以把string暂时理解为:
C++ 标准库封装好的“可自动管理空间的字符串类”。
它不仅保存字符串本身,还提供了大量成员函数:
s.size(); s += "world"; s.find("or"); s.substr(...); s.clear();所以学习string时,实际上要掌握两个层次:
第一层:会使用 string ↓ 构造、访问、遍历、增删、查找、截取、输入输出 第二层:理解 string 为什么能这样工作 ↓ 动态内存、构造函数、析构函数 ↓ 浅拷贝 / 深拷贝 ↓ 拷贝构造 / operator=2.string的构造:怎么创建一个字符串对象?
有几个常见构造方式
string s1;创建空字符串:
s1 = ""可以使用 C 风格字符串构造:
string s2("hello");也可以:
string s2 = "hello";还可以创建 n 个相同字符:
string s3(5, 'A');结果:
AAAAA还可以进行拷贝构造:
string s4(s2);相当于:
s2 = "hello" s4 = "hello"因此最常见的四种情况就是:
string s1; // 空字符串 string s2("hello"); // C字符串构造 string s3(5, 'x'); // xxxxx string s4(s2); // 拷贝构造二:auto和范围for
1.auto
auto的核心含义是:
让编译器根据初始化表达式自动推导变量类型。
例如:
int a = 10; auto b = a; auto c = 'A';编译器实际上会推导成:
int b = a; char c = 'A';特别适合类型很长的情况。
这里要记住一个很重要的区别:
int x = 10; auto a = x; auto& b = x;a是一个新的变量:
a ──> 10 x ──> 10修改a不影响x。
而:
auto& b = x;b是x的引用:
修改b就是在修改x
2.范围for
传统遍历:
string s = "hello"; for (int i = 0; i < s.size(); i++) { cout << s[i]; }C++11 可以写成:
for (auto ch : s) { cout << ch; }范围for可以用于数组和容器对象,它会自动完成迭代、取数据和结束判断;对于容器,底层遍历思想可以理解为利用迭代器完成
这里尤其重要的是:
for (auto ch : s)和:
for (auto& ch : s)是不一样的。
前者:
for (auto ch : s) { ch = 'A'; }ch是字符的副本,所以不会真正修改字符串。
而:
for (auto& ch : s) { ch = 'A'; }ch是原字符的引用,可以直接修改string。
比如可以对字符进行乘法操作
for (auto& e : array) { e *= 2; }三:string类的常用接口
(一)string类对象的容量操作
1.size(),length(),capacity()
size()和length()都是返回字符串有效字符长度
capacity()是返回空间总大小
他们的返回值都是size_t类型的
理解它们之前,一定要区分:
size = 当前有多少个有效字符 capacity = 当前已经准备了多少存储空间例如:
string s = "hello";逻辑上:
有效字符: h e l l o ↑ ↑ 共 5 个 size() = 5而capacity()可能比 5 大,因为string为了以后追加字符,可能提前准备额外空间。
可以类比成:
宿舍当前住 5 人 → size = 5 宿舍最多能住 10 人 → capacity = 10例如:
两者底层作用相同,通常更常使用size(),因为它和其他容器的接口保持一致
2.empty()
empty()是检测字符串释放为空串,是返回true,否则返回false
它的返回值类型是bool类型
string s2 = ""; if (s2.empty()) { cout << "字符是空" << endl; } else cout << "字符不为空" << endl;相当于判断:
s.size() == 0但:
s.empty()可读性更好。
3.clear()
clear()只清字符,不一定释放空间
例如:
string s = "hello"; s.clear();此时:
s.size() == 0但是特别强调:
clear()只是清除有效字符,并不会因此改变底层已经拥有的空间大小
可以理解为:
原来: size = 5 capacity = 15 clear以后: size = 0 capacity = 15宿舍的人走了,但房间没有拆
4.reserve()
reserve():提前准备空间
例如你准备往字符串里放很多字符:
string s; s.reserve(100);意思不是:字符串现在有100个字符
而是:提前准备能够容纳大约100个字符的空间
所以:
s.size()仍然是:
0注意:
reserve改变的是预留空间,不改变有效元素个数
它的用途主要是提高效率。
假如你不断:
s += 'a'; s += 'b'; s += 'c'; ...空间不够时可能需要:
申请新空间 ↓ 复制旧数据 ↓ 释放旧空间如果提前知道大概需要 1000 个字符:
s.reserve(1000);就可以减少重新申请空间的次数。
所以如果能够预估字符串大概会存放多少字符,可以提前使用reserve。
5.resize()
resize():将有效字符的个数该成n个,多出的空间用字符c填充resize()真正改变size(),它与reserve()不同假设:
string s = "hello";执行:
s.resize(3);得到:
hel因为有效字符数量变成了 3。
如果:
s.resize(8, 'x');得到:
helloxxx当增加字符个数时:
resize(n)会用默认值填补新增位置,而:
resize(n, c)会用字符c填补;如果缩小字符串,则底层空间总大小通常不会因此缩小
所以你可以牢牢记住:
(二)string类对象的访问及遍历操作
1.operator:返回pos位置的字符
最常见的方式:
operator[]例如:
string s = "hello"; cout << s[0];输出:
h也可以修改:
s[0] = 'H';得到:
Hello可以把:
s[i]理解成:访问字符串中的第 i 个字符
下标从0开始
所以:
hello 012342.begin+end
首先我们要制动什么是迭代器呢?
你可以先把它理解成:
一种“类似指针”的对象,用来访问容器中的元素。
它的主要作用是:按照一定顺序遍历vector、string、list、map等容器中的元素,而不需要关心这些容器内部到底是怎么存储数据的
begin():指向第一个字符
例如:
string s = "hello"; auto it = s.begin(); cout << *it;输出:h
因为:
hello ↑ begin()也可以修改字符:
string s = "hello"; auto it = s.begin(); *it = 'H'; cout << s;输出:Hello
因为:
*it代表当前迭代器指向的字符
end():指向最后一个字符后面
这是最容易搞错的地方。
对于:
string s = "hello";不是:
h e l l o ↑ end()而是:
h e l l o [ ] ↑ end()所以不能直接:
cout << *s.end();这是错误的,因为end()不指向有效字符。
如果想通过end()得到最后一个字符,可以先往前移动一次:
string s = "hello"; auto it = s.end(); --it; cout << *it;输出:o
因为:
初始: h e l l o [ ] ↑ it 执行 --it 后: h e l l o [ ] ↑ it3.rbegin+rend
rbegin():反向遍历的起点
rbegin()可以理解成:
reverse begin,也就是反向遍历时的第一个位置。
对于:
string s = "hello";rbegin()指向最后一个字符:
h e l l o ↑ rbegin()例如:
string s = "hello"; auto it = s.rbegin(); cout << *it;输出:o
rend():反向遍历的结束位置
rend()是:
reverse end,反向遍历结束的位置。
它位于第一个字符之前:
[ ] h e l l o ↑ rend()所以反向遍历:
#include <iostream> #include <string> using namespace std; int main() { string s = "hello"; for (auto it = s.rbegin(); it != s.rend(); ++it) { cout << *it << " "; } return 0; }输出:
o l l e h注意这里一个比较有意思的地方:
++it;虽然写的是++,但是由于这是反向迭代器,所以实际方向是从右往左:
o → l → l → e → h也就是说:普通迭代器 ++是向右走。
而:反向迭代器 ++是向左走。
4.三种非常重要的遍历方式:
第一种是下标:
for (size_t i = 0; i < s.size(); ++i) { cout << s[i]; }第二种是迭代器:
for (auto it = s.begin(); it != s.end(); ++it) { cout << *it; }begin() end()用于获得遍历区间,另外还有:
rbegin() rend()进行反向遍历
第三种就是最方便的范围for:
for (auto ch : s) { cout << ch; }需要修改时:
for (auto& ch : s) { ch += 1; }(三)string类对象的修改操作
1.push_back、append、+=
push_back:在字符串后尾插字符c
append:在字符串后追加一个字符串
+=:在字符串后追加字符串str
例如:
string s = "hello"; s.push_back('!');得到:
hello!push_back主要添加一个字符。
而:
s.append(" world");得到:
hello world最常见的还是:
s += '!'; s += " world";+=既可以连接字符,也可以连接字符串,因此实际使用中很方便
2.c_str():
c_str()是std::string提供的一个成员函数,用来把 C++ 的string转成C 风格字符串,也就是const char*
比如:
#include <iostream> #include <string> using namespace std; int main() { string s = "hello"; const char* p = s.c_str(); cout << p << endl; return 0; }输出:
hello你可以把它理解为:
string s = "hello";内部保存的是字符串内容,而:
s.c_str()会返回一个指向字符序列的指针:
h e l l o \0 ↑ p最后的\0是 C 风格字符串的结束标志
所以c_str()最常见的用途,就是在一些只接受const char*的 C 函数或旧式接口里使用std::string。
例如printf:
#include <cstdio> #include <string> using namespace std; int main() { string s = "hello"; printf("%s\n", s.c_str()); return 0; }因为%s需要的是:
const char*而不是:
std::string所以不能直接这样写:
printf("%s", s); // 错误而要写:
printf("%s", s.c_str()); // 正确3.find和npos
find()用来查找字符或子串的位置;如果没找到,就返回string::npos
假设有:
string s = "hello world";查找字符:
size_t pos = s.find('o'); cout << pos;输出:4
因为字符串下标是从0开始的:
h e l l o w o r l d 0 1 2 3 4 5 6 7 8 9 10 ↑ o所以第一个'o'的位置是4。
如果查找一个不存在的字符:
size_t pos = s.find('x');这时候不会返回-1来表示失败,而是返回:
string::npos因此通常这样写:
string s = "hello world"; size_t pos = s.find('x'); if (pos == string::npos) { cout << "没有找到"; } else { cout << "找到了,位置是:" << pos; }输出:没有找到
find()也可以查找子字符串:
string s = "hello world"; size_t pos = s.find("world"); cout << pos;输出:6
因为:
hello world ↑ 6也就是"world"从下标6开始。
完整写法通常是:
string s = "hello world"; size_t pos = s.find("world"); if (pos != string::npos) { cout << "找到了" << endl; cout << "起始位置:" << pos << endl; } else { cout << "没有找到" << endl; }find()默认找第一个匹配位置
例如:
string s = "banana"; size_t pos = s.find('a'); cout << pos;输出:1
虽然banana中有很多个a:
b a n a n a 0 1 2 3 4 5 ↑ ↑ ↑但是:
s.find('a')只返回第一个:1
find可以指定从哪里开始找
例如:
string s = "banana"; size_t pos = s.find('a', 2); cout << pos;这里表示:
从下标
2开始寻找'a'。
字符串:
b a n a n a 0 1 2 3 4 5 ↑ 从这里开始所以找到的是:3,而不是1。
如何找到所有相同字符?
例如:
string s = "banana";我们想找到所有'a':
size_t pos = s.find('a'); while (pos != string::npos) { cout << pos << " "; pos = s.find('a', pos + 1); }输出:
1 3 54.rfind
rfind()是std::string中用来从后往前查找字符或子字符串的函数
例如:
string s = "banana";字符串下标是:
b a n a n a 0 1 2 3 4 5如果写:
cout << s.find('a');输出:1
因为find()找的是第一个'a'。
而:
cout << s.rfind('a');输出:5
因为rfind()从后往前找,所以找到的是最后一个'a'
虽然rfind()是从右往左查找,但是它返回的仍然是字符串正常的下标
查找子字符串
rfind()不仅能查一个字符,也可以查字符串。
例如:
string s = "abcabcabc"; size_t pos = s.rfind("abc"); cout << pos;输出:6
因为:
a b c a b c a b c 0 1 2 3 4 5 6 7 8 ↑ abc"abc"一共出现了三次:
abc abc abc ↑ ↑ ↑ 0 3 6rfind()找最后一次出现的位置:
6而:
s.find("abc")返回:
0所以区别非常直观:
s.find("abc"); // 第一次出现的位置 s.rfind("abc"); // 最后一次出现的位置找不到时仍然返回string::npos
这和find()完全一样。
例如:
string s = "hello"; size_t pos = s.rfind('x'); if (pos == string::npos) { cout << "没有找到"; } else { cout << "找到了:" << pos; }输出:没有找到
所以标准写法依然是:
if (s.rfind("abc") != string::npos) { cout << "找到了"; }可以指定从哪个位置往前找
它还有一种形式:
s.rfind(要找的内容, 起始位置);不过这里要特别注意:
rfind()的第二个参数表示:从这个下标开始,向前查找。
例如:
string s = "banana"; size_t pos = s.rfind('a', 4); cout << pos;字符串:
b a n a n a 0 1 2 3 4 5 ↑ 从这里开始从下标4开始向左:
4 → 3 → 2 → 1 → 0首先遇到'a'的位置是:3
所以输出:3
5.substr
语法可以理解为:
s.substr(pos, n);即:
从
pos位置开始,取n个字符。
定义为:从pos开始截取n个字符并返回
例如:
string s = "hello world"; string t = s.substr(6, 5);得到:
world所以:
hello world ↑ pos = 6取 5 个:
worldfind + substr经常一起出现:
size_t pos = s.find(':'); string left = s.substr(0, pos); string right = s.substr(pos + 1);