\r\n
\r\n
\r\n
背景
\r\n
早期,阿里巴巴B2B公司因为存在杭州和美国双机房部署,存在跨机房同步的业务需求。不过早期的数据库同步业务,主要是基于trigger的方式获取增量变更,不过从2010年开始,阿里系公司开始逐步的尝试基于数据库的日志解析,获取增量变更进行同步,由此衍生出了增量订阅&消费的业务,从此开启了一段新纪元。ps. 目前内部使用的同步,已经支持mysql5.x和oracle部分版本的日志解析
\r\n
\r\n
基于日志增量订阅&消费支持的业务:
\r\n
\r\n- 数据库镜像
\r\n- 数据库实时备份
\r\n- 多级索引 (卖家和买家各自分库索引)
\r\n- search build
\r\n- 业务cache刷新
\r\n- 价格变化等重要业务消息
\r\n
\r\n
项目介绍
\r\n
名称:canal [kə'næl]
\r\n
译意: 水道/管道/沟渠
\r\n
语言: 纯java开发
\r\n
定位: 基于数据库增量日志解析,提供增量数据订阅&消费,目前主要支持了mysql
\r\n
\r\n
工作原理
\r\n
mysql主备复制实现
\r\n

从上层来看,复制分成三步:
\r\n
\r\n- master将改变记录到二进制日志(binary log)中(这些记录叫做二进制日志事件,binary log events,可以通过show binlog events进行查看);
\r\n- slave将master的binary log events拷贝到它的中继日志(relay log);
\r\n- slave重做中继日志中的事件,将改变反映它自己的数据。
\r\n
\r\n
canal的工作原理:
\r\n

\r\n
原理相对比较简单:
\r\n
\r\n- canal模拟mysql slave的交互协议,伪装自己为mysql slave,向mysql master发送dump协议
\r\n- mysql master收到dump请求,开始推送binary log给slave(也就是canal)
\r\n- canal解析binary log对象(原始为byte流)
\r\n
\r\n
架构
\r\n

\r\n
说明:
\r\n
\r\n- server代表一个canal运行实例,对应于一个jvm
\r\n- instance对应于一个数据队列 (1个server对应1..n个instance)
\r\n
\r\n
instance模块:
\r\n
\r\n- eventParser (数据源接入,模拟slave协议和master进行交互,协议解析)
\r\n- eventSink (Parser和Store链接器,进行数据过滤,加工,分发的工作)
\r\n- eventStore (数据存储)
\r\n- metaManager (增量订阅&消费信息管理器)
\r\n
\r\n
\r\n
Entry\r\n Header\r\n\t\tlogfileName [binlog文件名]\r\n\t\tlogfileOffset [binlog position]\r\n\t\texecuteTime [发生的变更]\r\n\t\tschemaName \r\n\t\ttableName\r\n\t\teventType [insert/update/delete类型]\r\n\tentryType \t[事务头BEGIN/事务尾END/数据ROWDATA]\r\n\tstoreValue \t[byte数据,可展开,对应的类型为RowChange]\r\n\t\r\nRowChange\r\n\tisDdl\t\t[是否是ddl变更操作,比如create table/drop table]\r\n\tsql\t\t[具体的ddl sql]\r\n\trowDatas\t[具体insert/update/delete的变更数据,可为多条,1个binlog event事件可对应多条变更,比如批处理]\r\n\t\tbeforeColumns [Column类型的数组]\r\n\t\tafterColumns [Column类型的数组]\r\n\t\t\r\nColumn \r\n\tindex\t\t\r\n\tsqlType\t\t[jdbc type]\r\n\tname\t\t[column name]\r\n\tisKey\t\t[是否为主键]\r\n\tupdated\t\t[是否发生过变更]\r\n\tisNull\t\t[值是否为null]\r\n\tvalue\t\t[具体的内容,注意为文本]
\r\n
说明:
\r\n
\r\n- 可以提供数据库变更前和变更后的字段内容,针对binlog中没有的name,isKey等信息进行补全
\r\n- 可以提供ddl的变更语句
\r\n
\r\n
QuickStart
\r\n
几点说明:(mysql初始化)
\r\n
a. canal的原理是基于mysql binlog技术,所以这里一定需要开启mysql的binlog写入功能,并且配置binlog模式为row.
\r\n
[mysqld]\r\nlog-bin=mysql-bin #添加这一行就ok\r\nbinlog-format=ROW #选择row模式\r\nserver_id=1 #配置mysql replaction需要定义,不能和canal的slaveId重复
\r\nb. canal的原理是模拟自己为mysql slave,所以这里一定需要做为mysql slave的相关权限.
\r\n
\r\n
CREATE USER canal IDENTIFIED BY 'canal'; \r\nGRANT SELECT, REPLICATION SLAVE, REPLICATION CLIENT ON *.* TO 'canal'@'%';\r\n-- GRANT ALL PRIVILEGES ON *.* TO 'canal'@'%' ;\r\nFLUSH PRIVILEGES;
\r\n
针对已有的账户可通过grants查询权限:
\r\n
启动步骤:
\r\n
1. 下载canal
\r\n
下载部署包
\r\n
wget http://canal4mysql.googlecode.com/files/canal.deployer-1.0.0.tar.gz
\r\n
or
\r\n
自己编译
\r\n
git clone git@github.com:otter-projects/canal.git\r\ncd canal; \r\nmvn clean install -Dmaven.test.skip -Denv=release
\r\n
编译完成后,会在根目录下产生target/canal.deployer-$version.tar.gz
\r\n
\r\n
2. 解压缩
\r\n
mkdir /tmp/canal\r\ntar zxvf canal.deployer-1.0.0.tar.gz -C /tmp/canal
\r\n
\r\n
解压完成后,进入/tmp/canal目录,可以看到如下结构:
\r\n
\r\n
drwxr-xr-x 2 jianghang jianghang 136 2013-02-05 21:51 bin\r\ndrwxr-xr-x 4 jianghang jianghang 160 2013-02-05 21:51 conf\r\ndrwxr-xr-x 2 jianghang jianghang 1.3K 2013-02-05 21:51 lib\r\ndrwxr-xr-x 2 jianghang jianghang 48 2013-02-05 21:29 logs
\r\n
\r\n
3. 配置修改
\r\n
\r\n
公用参数:
\r\n
vi conf/canal.properties
\r\n
#################################################\r\n######### common argument ############# \r\n#################################################\r\ncanal.id= 1\r\ncanal.address=\r\ncanal.port= 11111\r\ncanal.zkServers=\r\n# flush data to zk\r\ncanal.zookeeper.flush.period = 1000\r\n## memory store RingBuffer size, should be Math.pow(2,n)\r\ncanal.instance.memory.buffer.size = 32768\r\n\r\n## detecing config\r\ncanal.instance.detecting.enable = false\r\ncanal.instance.detecting.sql = insert into retl.xdual values(1,now()) on duplicate key update x=now()\r\ncanal.instance.detecting.interval.time = 3 \r\ncanal.instance.detecting.retry.threshold = 3 \r\ncanal.instance.detecting.heartbeatHaEnable = false\r\n\r\n# support maximum transaction size, more than the size of the transaction will be cut into multiple transactions delivery\r\ncanal.instance.transactionn.size = 1024\r\n\r\n# network config\r\ncanal.instance.network.receiveBufferSize = 16384\r\ncanal.instance.network.sendBufferSize = 16384\r\ncanal.instance.network.soTimeout = 30\r\n\r\n#################################################\r\n######### destinations ############# \r\n#################################################\r\ncanal.destinations= example\r\n\r\ncanal.instance.global.mode = spring \r\ncanal.instance.global.lazy = true ##修改为false,代表立马启动\r\n#canal.instance.global.manager.address = 127.0.0.1:1099\r\ncanal.instance.global.spring.xml = classpath:spring/memory-instance.xml\r\n#canal.instance.global.spring.xml = classpath:spring/default-instance.xml
\r\n
\r\n
应用参数:
\r\n
vi conf/example/instance.properties
\r\n
#################################################\r\n## mysql serverId\r\ncanal.instance.mysql.slaveId = 1234\r\n\r\n# position info\r\ncanal.instance.master.address = 127.0.0.1:3306 #改成自己的数据库地址\r\ncanal.instance.master.journal.name = \r\ncanal.instance.master.position = \r\ncanal.instance.master.timestamp = \r\n\r\n#canal.instance.standby.address = \r\n#canal.instance.standby.journal.name =\r\n#canal.instance.standby.position = \r\n#canal.instance.standby.timestamp = \r\n\r\n# username/password\r\ncanal.instance.dbUsername = retl #改成自己的数据库信息\r\ncanal.instance.dbPassword = retl #改成自己的数据库信息\r\ncanal.instance.defaultDatabaseName = #改成自己的数据库信息\r\ncanal.instance.connectionCharsetNumber = 33 #改成自己的数据库信息\r\ncanal.instance.connectionCharset = UTF-8 #改成自己的数据库信息\r\n\r\n# table regex\r\ncanal.instance.filter.regex = .*\\\\..*\r\n\r\n#################################################\r\n
\r\n
\r\n
\r\n
说明:
\r\n
\r\n- canal.instance.connectionCharset 代表数据库的编码方式对应到java中的编码类型,比如UTF-8,GBK , ISO-8859-1
\r\n- canal.instance.connectionCharsetNumber 代表数据库的编码方式对应mysql中的唯一id,详细的映射关系可查看:com.mysql.jdbc.CharsetMapping.INDEX_TO_CHARSET
针对常见的编码:
utf-8 <=> 33
gb2312 <=> 24
gbk <=> 28 \r\n
\r\n
4. 准备启动
\r\n
\r\n
sh bin/startup.sh
\r\n
\r\n
5. 查看日志
\r\n
vi logs/canal/canal.log
\r\n
2013-02-05 22:45:27.967 [main] INFO com.alibaba.otter.canal.deployer.CanalLauncher - ## start the canal server.\r\n2013-02-05 22:45:28.113 [main] INFO com.alibaba.otter.canal.deployer.CanalController - ## start the canal server[10.1.29.120:11111]\r\n2013-02-05 22:45:28.210 [main] INFO com.alibaba.otter.canal.deployer.CanalLauncher - ## the canal server is running now ......
\r\n
\r\n
具体instance的日志:
\r\n
vi logs/example/example.log
\r\n
2013-02-05 22:50:45.636 [main] INFO c.a.o.c.i.spring.support.PropertyPlaceholderConfigurer - Loading properties file from class path resource [canal.properties]\r\n2013-02-05 22:50:45.641 [main] INFO c.a.o.c.i.spring.support.PropertyPlaceholderConfigurer - Loading properties file from class path resource [example/instance.properties]\r\n2013-02-05 22:50:45.803 [main] INFO c.a.otter.canal.instance.spring.CanalInstanceWithSpring - start CannalInstance for 1-example \r\n2013-02-05 22:50:45.810 [main] INFO c.a.otter.canal.instance.spring.CanalInstanceWithSpring - start successful....
\r\n
\r\n
6. 关闭
\r\n
sh bin/stop.sh
\r\n
\r\n
it's over.
\r\n
\r\n
ClientExample
\r\n
依赖配置:(目前暂未正式发布到mvn仓库,所以需要各位下载canal源码后手工执行下mvn clean install -Dmaven.test.skip)
\r\n
<dependency>\r\n <groupId>com.alibaba.otter</groupId>\r\n <artifactId>canal.client</artifactId>\r\n <version>1.0.0</version>\r\n</dependency>
\r\n
\r\n
1. 创建mvn标准工程:
\r\n
mvn archetype:create -DgroupId=com.alibaba.otter -DartifactId=canal.sample
\r\n
\r\n
2. 修改pom.xml,添加依赖
\r\n
\r\n
3. ClientSample代码
\r\n
package com.alibaba.otter.canal.sample;\r\n\r\nimport java.net.InetSocketAddress;\r\nimport java.util.List;\r\n\r\nimport com.alibaba.otter.canal.common.utils.AddressUtils;\r\nimport com.alibaba.otter.canal.protocol.Message;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.Column;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.Entry;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.EntryType;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.EventType;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.RowChange;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.RowData;\r\n\r\npublic class SimpleCanalClientExample {\r\n\r\n public static void main(String args[]) {\r\n // 创建链接\r\n CanalConnector connector = CanalConnectors.newSingleConnector(new InetSocketAddress(AddressUtils.getHostIp(),\r\n 11111), \"example\", \"\", \"\");\r\n int batchSize = 1000;\r\n int emptyCount = 0;\r\n try {\r\n connector.connect();\r\n connector.subscribe(\".*\\\\..*\");\r\n connector.rollback();\r\n int totalEmtryCount = 120;\r\n while (emptyCount < totalEmtryCount) {\r\n Message message = connector.getWithoutAck(batchSize); // 获取指定数量的数据\r\n long batchId = message.getId();\r\n int size = message.getEntries().size();\r\n if (batchId == -1 || size == 0) {\r\n emptyCount++;\r\n System.out.println(\"empty count : \" + emptyCount);\r\n try {\r\n Thread.sleep(1000);\r\n } catch (InterruptedException e) {\r\n }\r\n } else {\r\n emptyCount = 0;\r\n // System.out.printf(\"message[batchId=%s,size=%s] \\n\", batchId, size);\r\n printEntry(message.getEntries());\r\n }\r\n\r\n connector.ack(batchId); // 提交确认\r\n // connector.rollback(batchId); // 处理失败, 回滚数据\r\n }\r\n\r\n System.out.println(\"empty too many times, exit\");\r\n } finally {\r\n connector.disconnect();\r\n }\r\n }\r\n\r\n private static void printEntry(List<Entry> entrys) {\r\n for (Entry entry : entrys) {\r\n if (entry.getEntryType() == EntryType.TRANSACTIONBEGIN || entry.getEntryType() == EntryType.TRANSACTIONEND) {\r\n continue;\r\n }\r\n\r\n RowChange rowChage = null;\r\n try {\r\n rowChage = RowChange.parseFrom(entry.getStoreValue());\r\n } catch (Exception e) {\r\n throw new RuntimeException(\"ERROR ## parser of eromanga-event has an error , data:\" + entry.toString(),\r\n e);\r\n }\r\n\r\n EventType eventType = rowChage.getEventType();\r\n System.out.println(String.format(\"================> binlog[%s:%s] , name[%s,%s] , eventType : %s\",\r\n entry.getHeader().getLogfileName(), entry.getHeader().getLogfileOffset(),\r\n entry.getHeader().getSchemaName(), entry.getHeader().getTableName(),\r\n eventType));\r\n\r\n for (RowData rowData : rowChage.getRowDatasList()) {\r\n if (eventType == EventType.DELETE) {\r\n printColumn(rowData.getBeforeColumnsList());\r\n } else if (eventType == EventType.INSERT) {\r\n printColumn(rowData.getAfterColumnsList());\r\n } else {\r\n System.out.println(\"-------> before\");\r\n printColumn(rowData.getBeforeColumnsList());\r\n System.out.println(\"-------> after\");\r\n printColumn(rowData.getAfterColumnsList());\r\n }\r\n }\r\n }\r\n }\r\n\r\n private static void printColumn(List<Column> columns) {\r\n for (Column column : columns) {\r\n System.out.println(column.getName() + \" : \" + column.getValue() + \" update=\" + column.getUpdated());\r\n }\r\n }\r\n}
\r\n
\r\n
4. 运行Client
\r\n
首先启动Canal Server,可参加QuickStart : http://agapple.iteye.com/blogs/1796070
\r\n
启动Canal Client后,可以从控制台从看到类似消息:
\r\n
empty count : 1\r\nempty count : 2\r\nempty count : 3\r\nempty count : 4
\r\n
此时代表当前数据库无变更数据
\r\n
\r\n
5. 触发数据库变更
\r\n
mysql> use test;\r\nDatabase changed\r\nmysql> CREATE TABLE `xdual` (\r\n -> `ID` int(11) NOT NULL AUTO_INCREMENT,\r\n -> `X` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP,\r\n -> PRIMARY KEY (`ID`)\r\n -> ) ENGINE=InnoDB AUTO_INCREMENT=3 DEFAULT CHARSET=utf8 ;\r\nQuery OK, 0 rows affected (0.06 sec)\r\n\r\nmysql> insert into xdual(id,x) values(null,now());Query OK, 1 row affected (0.06 sec)
\r\n
\r\n
可以从控制台中看到:
\r\n
empty count : 1\r\nempty count : 2\r\nempty count : 3\r\nempty count : 4\r\n================> binlog[mysql-bin.001946:313661577] , name[test,xdual] , eventType : INSERT\r\nID : 4 update=true\r\nX : 2013-02-05 23:29:46 update=true
\r\n
\r\n
最后:
\r\n
整个代码在附件中可以下载,如有问题可及时联系。
\r\n
\r\n \r\n