{"name":"Canal","tagline":"阿里巴巴mysql数据库binlog的增量订阅&消费组件","body":"
\r\n
\r\n

\r\n

背景

\r\n

早期,阿里巴巴B2B公司因为存在杭州和美国双机房部署,存在跨机房同步的业务需求。不过早期的数据库同步业务,主要是基于trigger的方式获取增量变更,不过从2010年开始,阿里系公司开始逐步的尝试基于数据库的日志解析,获取增量变更进行同步,由此衍生出了增量订阅&消费的业务,从此开启了一段新纪元。ps. 目前内部使用的同步,已经支持mysql5.x和oracle部分版本的日志解析

\r\n

\r\n

基于日志增量订阅&消费支持的业务:

\r\n
    \r\n
  1. 数据库镜像
  2. \r\n
  3. 数据库实时备份
  4. \r\n
  5. 多级索引 (卖家和买家各自分库索引)
  6. \r\n
  7. search build
  8. \r\n
  9. 业务cache刷新
  10. \r\n
  11. 价格变化等重要业务消息
  12. \r\n
\r\n

项目介绍

\r\n

名称:canal [kə'næl]

\r\n

译意: 水道/管道/沟渠

\r\n

语言: 纯java开发

\r\n

定位: 基于数据库增量日志解析,提供增量数据订阅&消费,目前主要支持了mysql

\r\n

\r\n

工作原理

\r\n

mysql主备复制实现

\r\n

\"\"
从上层来看,复制分成三步:

\r\n
    \r\n
  1. master将改变记录到二进制日志(binary log)中(这些记录叫做二进制日志事件,binary log events,可以通过show binlog events进行查看);
  2. \r\n
  3. slave将master的binary log events拷贝到它的中继日志(relay log);
  4. \r\n
  5. slave重做中继日志中的事件,将改变反映它自己的数据。
  6. \r\n
\r\n

canal的工作原理:

\r\n

\"\"

\r\n

原理相对比较简单:

\r\n
    \r\n
  1. canal模拟mysql slave的交互协议,伪装自己为mysql slave,向mysql master发送dump协议
  2. \r\n
  3. mysql master收到dump请求,开始推送binary log给slave(也就是canal)
  4. \r\n
  5. canal解析binary log对象(原始为byte流)
  6. \r\n
\r\n

架构

\r\n

\"\"

\r\n

说明:

\r\n\r\n

instance模块:

\r\n\r\n

数据对象格式:EntryProtocol.proto\r\n

\r\n
Entry\r\n    Header\r\n\t\tlogfileName [binlog文件名]\r\n\t\tlogfileOffset [binlog position]\r\n\t\texecuteTime [发生的变更]\r\n\t\tschemaName \r\n\t\ttableName\r\n\t\teventType [insert/update/delete类型]\r\n\tentryType \t[事务头BEGIN/事务尾END/数据ROWDATA]\r\n\tstoreValue \t[byte数据,可展开,对应的类型为RowChange]\r\n\t\r\nRowChange\r\n\tisDdl\t\t[是否是ddl变更操作,比如create table/drop table]\r\n\tsql\t\t[具体的ddl sql]\r\n\trowDatas\t[具体insert/update/delete的变更数据,可为多条,1个binlog event事件可对应多条变更,比如批处理]\r\n\t\tbeforeColumns [Column类型的数组]\r\n\t\tafterColumns [Column类型的数组]\r\n\t\t\r\nColumn \r\n\tindex\t\t\r\n\tsqlType\t\t[jdbc type]\r\n\tname\t\t[column name]\r\n\tisKey\t\t[是否为主键]\r\n\tupdated\t\t[是否发生过变更]\r\n\tisNull\t\t[值是否为null]\r\n\tvalue\t\t[具体的内容,注意为文本]
\r\n

说明:

\r\n\r\n

QuickStart

\r\n

几点说明:(mysql初始化)

\r\n

a. canal的原理是基于mysql binlog技术,所以这里一定需要开启mysql的binlog写入功能,并且配置binlog模式为row.

\r\n
[mysqld]\r\nlog-bin=mysql-bin #添加这一行就ok\r\nbinlog-format=ROW #选择row模式\r\nserver_id=1 #配置mysql replaction需要定义,不能和canal的slaveId重复
\r\nb. canal的原理是模拟自己为mysql slave,所以这里一定需要做为mysql slave的相关权限.
\r\n
\r\n
CREATE USER canal IDENTIFIED BY 'canal';  \r\nGRANT SELECT, REPLICATION SLAVE, REPLICATION CLIENT ON *.* TO 'canal'@'%';\r\n-- GRANT ALL PRIVILEGES ON *.* TO 'canal'@'%' ;\r\nFLUSH PRIVILEGES;
\r\n

针对已有的账户可通过grants查询权限:

\r\n

启动步骤:

\r\n

1. 下载canal

\r\n

下载部署包

\r\n
wget http://canal4mysql.googlecode.com/files/canal.deployer-1.0.0.tar.gz
\r\n

or

\r\n

自己编译

\r\n
git clone git@github.com:otter-projects/canal.git\r\ncd canal; \r\nmvn clean install -Dmaven.test.skip -Denv=release
\r\n

编译完成后,会在根目录下产生target/canal.deployer-$version.tar.gz

\r\n

\r\n

2. 解压缩

\r\n
mkdir /tmp/canal\r\ntar zxvf canal.deployer-1.0.0.tar.gz  -C /tmp/canal
\r\n

\r\n

解压完成后,进入/tmp/canal目录,可以看到如下结构:

\r\n

\r\n
drwxr-xr-x 2 jianghang jianghang  136 2013-02-05 21:51 bin\r\ndrwxr-xr-x 4 jianghang jianghang  160 2013-02-05 21:51 conf\r\ndrwxr-xr-x 2 jianghang jianghang 1.3K 2013-02-05 21:51 lib\r\ndrwxr-xr-x 2 jianghang jianghang   48 2013-02-05 21:29 logs
\r\n

\r\n

3. 配置修改

\r\n

\r\n

公用参数:

\r\n
vi conf/canal.properties
\r\n
#################################################\r\n#########               common argument         ############# \r\n#################################################\r\ncanal.id= 1\r\ncanal.address=\r\ncanal.port= 11111\r\ncanal.zkServers=\r\n# flush data to zk\r\ncanal.zookeeper.flush.period = 1000\r\n## memory store RingBuffer size, should be Math.pow(2,n)\r\ncanal.instance.memory.buffer.size = 32768\r\n\r\n## detecing config\r\ncanal.instance.detecting.enable = false\r\ncanal.instance.detecting.sql = insert into retl.xdual values(1,now()) on duplicate key update x=now()\r\ncanal.instance.detecting.interval.time = 3 \r\ncanal.instance.detecting.retry.threshold = 3 \r\ncanal.instance.detecting.heartbeatHaEnable = false\r\n\r\n# support maximum transaction size, more than the size of the transaction will be cut into multiple transactions delivery\r\ncanal.instance.transactionn.size =  1024\r\n\r\n# network config\r\ncanal.instance.network.receiveBufferSize = 16384\r\ncanal.instance.network.sendBufferSize = 16384\r\ncanal.instance.network.soTimeout = 30\r\n\r\n#################################################\r\n#########               destinations            ############# \r\n#################################################\r\ncanal.destinations= example\r\n\r\ncanal.instance.global.mode = spring \r\ncanal.instance.global.lazy = true  ##修改为false,代表立马启动\r\n#canal.instance.global.manager.address = 127.0.0.1:1099\r\ncanal.instance.global.spring.xml = classpath:spring/memory-instance.xml\r\n#canal.instance.global.spring.xml = classpath:spring/default-instance.xml
\r\n

\r\n

应用参数:

\r\n
vi conf/example/instance.properties
\r\n
#################################################\r\n## mysql serverId\r\ncanal.instance.mysql.slaveId = 1234\r\n\r\n# position info\r\ncanal.instance.master.address = 127.0.0.1:3306 #改成自己的数据库地址\r\ncanal.instance.master.journal.name = \r\ncanal.instance.master.position = \r\ncanal.instance.master.timestamp = \r\n\r\n#canal.instance.standby.address = \r\n#canal.instance.standby.journal.name =\r\n#canal.instance.standby.position = \r\n#canal.instance.standby.timestamp = \r\n\r\n# username/password\r\ncanal.instance.dbUsername = retl  #改成自己的数据库信息\r\ncanal.instance.dbPassword = retl  #改成自己的数据库信息\r\ncanal.instance.defaultDatabaseName =   #改成自己的数据库信息\r\ncanal.instance.connectionCharsetNumber = 33  #改成自己的数据库信息\r\ncanal.instance.connectionCharset = UTF-8  #改成自己的数据库信息\r\n\r\n# table regex\r\ncanal.instance.filter.regex = .*\\\\..*\r\n\r\n#################################################\r\n
\r\n

\r\n

\r\n

说明:

\r\n\r\n

4. 准备启动

\r\n

\r\n
sh bin/startup.sh
\r\n

\r\n

5. 查看日志

\r\n
vi logs/canal/canal.log
\r\n
2013-02-05 22:45:27.967 [main] INFO  com.alibaba.otter.canal.deployer.CanalLauncher - ## start the canal server.\r\n2013-02-05 22:45:28.113 [main] INFO  com.alibaba.otter.canal.deployer.CanalController - ## start the canal server[10.1.29.120:11111]\r\n2013-02-05 22:45:28.210 [main] INFO  com.alibaba.otter.canal.deployer.CanalLauncher - ## the canal server is running now ......
\r\n

\r\n

具体instance的日志:

\r\n
vi logs/example/example.log
\r\n
2013-02-05 22:50:45.636 [main] INFO  c.a.o.c.i.spring.support.PropertyPlaceholderConfigurer - Loading properties file from class path resource [canal.properties]\r\n2013-02-05 22:50:45.641 [main] INFO  c.a.o.c.i.spring.support.PropertyPlaceholderConfigurer - Loading properties file from class path resource [example/instance.properties]\r\n2013-02-05 22:50:45.803 [main] INFO  c.a.otter.canal.instance.spring.CanalInstanceWithSpring - start CannalInstance for 1-example \r\n2013-02-05 22:50:45.810 [main] INFO  c.a.otter.canal.instance.spring.CanalInstanceWithSpring - start successful....
\r\n

\r\n

6. 关闭

\r\n
sh bin/stop.sh
\r\n

\r\n

it's over.

\r\n
\r\n

ClientExample

\r\n

依赖配置:(目前暂未正式发布到mvn仓库,所以需要各位下载canal源码后手工执行下mvn clean install -Dmaven.test.skip)

\r\n
<dependency>\r\n    <groupId>com.alibaba.otter</groupId>\r\n    <artifactId>canal.client</artifactId>\r\n    <version>1.0.0</version>\r\n</dependency>
\r\n

\r\n

1. 创建mvn标准工程:

\r\n
mvn archetype:create -DgroupId=com.alibaba.otter -DartifactId=canal.sample
\r\n

\r\n

2. 修改pom.xml,添加依赖

\r\n

\r\n

3. ClientSample代码

\r\n
package com.alibaba.otter.canal.sample;\r\n\r\nimport java.net.InetSocketAddress;\r\nimport java.util.List;\r\n\r\nimport com.alibaba.otter.canal.common.utils.AddressUtils;\r\nimport com.alibaba.otter.canal.protocol.Message;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.Column;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.Entry;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.EntryType;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.EventType;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.RowChange;\r\nimport com.alibaba.otter.canal.protocol.CanalEntry.RowData;\r\n\r\npublic class SimpleCanalClientExample {\r\n\r\n    public static void main(String args[]) {\r\n        // 创建链接\r\n        CanalConnector connector = CanalConnectors.newSingleConnector(new InetSocketAddress(AddressUtils.getHostIp(),\r\n                                                                                            11111), \"example\", \"\", \"\");\r\n        int batchSize = 1000;\r\n        int emptyCount = 0;\r\n        try {\r\n            connector.connect();\r\n            connector.subscribe(\".*\\\\..*\");\r\n            connector.rollback();\r\n            int totalEmtryCount = 120;\r\n            while (emptyCount < totalEmtryCount) {\r\n                Message message = connector.getWithoutAck(batchSize); // 获取指定数量的数据\r\n                long batchId = message.getId();\r\n                int size = message.getEntries().size();\r\n                if (batchId == -1 || size == 0) {\r\n                    emptyCount++;\r\n                    System.out.println(\"empty count : \" + emptyCount);\r\n                    try {\r\n                        Thread.sleep(1000);\r\n                    } catch (InterruptedException e) {\r\n                    }\r\n                } else {\r\n                    emptyCount = 0;\r\n                    // System.out.printf(\"message[batchId=%s,size=%s] \\n\", batchId, size);\r\n                    printEntry(message.getEntries());\r\n                }\r\n\r\n                connector.ack(batchId); // 提交确认\r\n                // connector.rollback(batchId); // 处理失败, 回滚数据\r\n            }\r\n\r\n            System.out.println(\"empty too many times, exit\");\r\n        } finally {\r\n            connector.disconnect();\r\n        }\r\n    }\r\n\r\n    private static void printEntry(List<Entry> entrys) {\r\n        for (Entry entry : entrys) {\r\n            if (entry.getEntryType() == EntryType.TRANSACTIONBEGIN || entry.getEntryType() == EntryType.TRANSACTIONEND) {\r\n                continue;\r\n            }\r\n\r\n            RowChange rowChage = null;\r\n            try {\r\n                rowChage = RowChange.parseFrom(entry.getStoreValue());\r\n            } catch (Exception e) {\r\n                throw new RuntimeException(\"ERROR ## parser of eromanga-event has an error , data:\" + entry.toString(),\r\n                                           e);\r\n            }\r\n\r\n            EventType eventType = rowChage.getEventType();\r\n            System.out.println(String.format(\"================> binlog[%s:%s] , name[%s,%s] , eventType : %s\",\r\n                                             entry.getHeader().getLogfileName(), entry.getHeader().getLogfileOffset(),\r\n                                             entry.getHeader().getSchemaName(), entry.getHeader().getTableName(),\r\n                                             eventType));\r\n\r\n            for (RowData rowData : rowChage.getRowDatasList()) {\r\n                if (eventType == EventType.DELETE) {\r\n                    printColumn(rowData.getBeforeColumnsList());\r\n                } else if (eventType == EventType.INSERT) {\r\n                    printColumn(rowData.getAfterColumnsList());\r\n                } else {\r\n                    System.out.println(\"-------> before\");\r\n                    printColumn(rowData.getBeforeColumnsList());\r\n                    System.out.println(\"-------> after\");\r\n                    printColumn(rowData.getAfterColumnsList());\r\n                }\r\n            }\r\n        }\r\n    }\r\n\r\n    private static void printColumn(List<Column> columns) {\r\n        for (Column column : columns) {\r\n            System.out.println(column.getName() + \" : \" + column.getValue() + \"    update=\" + column.getUpdated());\r\n        }\r\n    }\r\n}
\r\n

\r\n

4. 运行Client

\r\n

首先启动Canal Server,可参加QuickStart : http://agapple.iteye.com/blogs/1796070

\r\n

启动Canal Client后,可以从控制台从看到类似消息:

\r\n
empty count : 1\r\nempty count : 2\r\nempty count : 3\r\nempty count : 4
\r\n

此时代表当前数据库无变更数据

\r\n

\r\n

5. 触发数据库变更

\r\n
mysql> use test;\r\nDatabase changed\r\nmysql> CREATE TABLE `xdual` (\r\n    ->   `ID` int(11) NOT NULL AUTO_INCREMENT,\r\n    ->   `X` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP,\r\n    ->   PRIMARY KEY (`ID`)\r\n    -> ) ENGINE=InnoDB AUTO_INCREMENT=3 DEFAULT CHARSET=utf8 ;\r\nQuery OK, 0 rows affected (0.06 sec)\r\n\r\nmysql> insert into xdual(id,x) values(null,now());Query OK, 1 row affected (0.06 sec)
\r\n

\r\n

可以从控制台中看到:

\r\n
empty count : 1\r\nempty count : 2\r\nempty count : 3\r\nempty count : 4\r\n================> binlog[mysql-bin.001946:313661577] , name[test,xdual] , eventType : INSERT\r\nID : 4    update=true\r\nX : 2013-02-05 23:29:46    update=true
\r\n

\r\n

最后:

\r\n

整个代码在附件中可以下载,如有问题可及时联系。

\r\n
\r\n \r\n
\r\ncanal.sample.tar.gz (2.2 KB)\r\n
\r\n","google":"UA-10379866-5","note":"Don't delete this file! It's used internally to help with page regeneration."}